Abstract
At the beginning of the twenty-first century, the primary view of infant visual attention development focused on a transition across the first postnatal year from being stimulus-driven to goal-driven, reflecting a broader shift from subcortical to cortical control. This perspective was supported by decades of infant looking-time studies. However, our understanding of infant attention has significantly evolved over the past 25 years, shaped by both theoretical advancements and new technological and methodological tools. Researchers now understand that attention development reflects multiple interacting systems that have cascading effects across time. The availability of infant-suitable eye-tracking methods have allowed researchers to consider multiple aspects of attention by precisely measuring when and where infants look, emerging quantitative models of stimulus saliency and computational models of the visual system have deepened our understanding of bottom-up and top-down influences on infant attention, and new methods to evaluate infants’ egocentric views have allowed researchers to measure attention in naturalistic contexts. Thus, these innovations allowed us to address questions that were unthinkable 25 years ago. Here, we discuss how these advances have transformed our understanding of infant attention development and outline key directions for future research, paving the way for even more exciting discoveries in the next 25 years.
Keywords: Visual attention, Attention development, Eye-tracking, Egocentric views, Computational tools
1. Introduction
Early philosophers emphasized the profound role of sensory experiences in shaping our understanding of the world, with Aristotle famously stating, “perception is the first form of knowledge.” This idea fueled a long line of scientific pursuits aimed to relate visual inputs, filtered by selective attention, to the acquisition of knowledge. Visual attention, or the ability to focus on one part of the visual environment while concurrently ignoring others, has cascading consequences for how people learn and interpret information (Posner, 1980; Treisman, 1964). It wasn’t until the late 1950s and early 1960s when researchers developed procedures to assess infants’ visual attention by measuring where and how long they looked, allowing us to draw conclusions about the perceptual and cognitive capabilities of infants (Fantz, 1956). It is important to point out that we now understand visual attention is not synonymous with looking; adult attention has often been studied in paradigms in which participants are explicitly instructed to not move their eyes (e.g., Posner cueing task; Posner, 1980). However, because infants cannot be given these verbal instructions, the study of visual attention in infancy has been dominated by measures of overt looking behaviors, and indeed, the terms “visual attention” and “looking” were used interchangeably in the first decades of research in this area (e.g., Cohen, 1973; Greenberg, 1977; Ruff & Turkewitz, 1975). The more recent introduction of physiological measures, including heart rate and EEG, has allowed researchers to ask questions about less observable aspects of attention as well, including engagement during sustained attention and covert attentional shifts (e.g., Colombo et al., 2001; Xie & Richards, 2016).
The primary conclusion drawn from the decades of research leading up to the end of the 20th century was that infants’ visual attention developed systematically across the first postnatal year. A large body of research focused on documenting and understanding the shift from reflexive to voluntary attention (e.g., Atkinson et al., 1996; Bronson, 1974, 1994), including how infants become able to inhibit reflexive attentional shifts (Atkinson et al., 1989; Johnson, 1995). This research also highlighted how very young infants’ attention is highly determined by bottom-up stimulus properties—such as high intensity contrast or salient movement—but increasingly with age, attention reflects top-down control based on higher-order factors, including familiarity and social significance (Colombo, 2001). These developments were attributed to a shift from subcortical to cortical processing (Bronson, 1974) as cortical regions involved in higher level cognitive functions gradually matured, such as those involved in goal-driven action and inhibition of prepotent responses (Johnson, 1990).
These initial conclusions were drawn primarily from studies in which human observers recorded infants’ looking times at one or two stimuli presented at a time (Fantz, 1956). These methods yielded valuable insights as they provide tight control over experimental stimuli. For instance, researchers explored how stimulus complexity, measured as the number of checks in a checkerboard, influenced infant attention while controlling for low-level perceptual factors (i.e., keeping the total number of black and white pixels constant across stimuli; Brennan et al., 1966). Although such methods reflected the technological limitations of the time, they laid a strong foundation for the field. Importantly, even as technological and theoretical advances have expanded the questions that can be addressed, researchers continue to rely on these fundamental methods for understanding the development of infant attention.
Over the first 25 years of the 21st century, our understanding of visual attention development has transformed due to significant technological and methodological advances in collecting and interpreting visual attention data from infants. Here, we discuss these advancements, highlighting the insights gained from these new findings. Our understanding of visual attention development has also become influenced by theoretical shifts in the field of developmental psychology that emphasize the role of interacting systems in the emergence and development of abilities (Oakes & Rakison, 2019; Smith & Thelen, 2003). As a result, researchers now consider attention development to reflect the interaction of multiple cognitive systems (Amso & Scerif, 2015; Amso & Tummeltshammer, 2020; Lynn et al., 2024; Reynolds & Romano, 2016) as well as the interaction between cognitive and other developing capacities (e.g., motor skills, perception) in the context of the whole child (Maia et al., 2024; Oakes, 2009, 2023a, 2023b; Smith & Gasser, 2005).
It is important to point out that much of the research on the development of visual attention has been largely conducted on infants from Western middle class families. According to sociocultural theory, children develop the ability to control their attention through their interactions with more experienced social partners within specific cultural and institutional contexts (e.g., Rogoff, 1990; Vygotsky, 1978). For instance, some cultures emphasize shared attention and guided exploration while others promote more independent attention (Bornstein et al., 1990; Bornstein et al., 1991). Although sociocultural theory has long shaped developmental thinking, there have been limited empirical comparisons of visual attention development across cultural contexts. This limited research suggests that basic attentional processes during infancy may indeed vary across cultures (e.g., DeBolt et al., 2025; Heise et al., 2025; Chavajay & Rogoff, 1999), socioeconomic statuses (e.g., Clearfield & Jedd, 2013), and language backgrounds (e.g., bilingual vs. monolingual; e.g., D’Souza et al., 2020; Singh et al., 2015). Thus conclusions from existing work must be considered in the cultural context in which they were observed, and assumptions about universality should be considered with caution.
2. Technological advances in the study of infant visual attention
In the last 25 years, a wide range of technological advancements has allowed for a much finer-grained understanding of the development of infant visual attention. For example, we now recognize multiple aspects of visual attention (e.g., overt versus covert orienting, sustained attention, disengagement) as a result of innovations in neuroimaging and physiological methods. This work often relies on multimethod approaches that combine behavioral measures of looking with neuroimaging measures, such as EEG (e.g., Pickron et al., 2018; Reynolds & Richards, 2005) or fNIRS (e.g., Emberson et al., 2017), as well as physiological measures such as heart rate deceleration (e.g., Colombo et al., 2001; Xie & Richards, 2016; Zhang et al., 2025). However, here we focus specifically on advances in the measurement and interpretation of infants’ looking behavior, building on the foundation laid by the decades of research beginning in the 1950s. In particular, we discuss how advances in 1) automated eye-tracking and 2) computational tools for quantifying stimulus properties have transformed the ways in which looking behaviors can be evaluated and related to specific stimulus properties (e.g., complexity, salience).
2.1. Automated eye-tracking
One of the most significant advances in the understanding of the development of visual attention was the introduction of automatic eye-tracking equipment and algorithms designed for collecting data from infants (Aslin & McMurray, 2004; Oakes, 2012). The types of manually coded data available prior to the 21st century (e.g., total looking time to a portion of a screen) became computable in a fraction of the time, facilitating research conducted using traditional looking time measures. These tools also enabled precise measurement of the spatial and temporal aspects of infants’ visual attention, revealing aspects of attention that traditional methods could not capture. Although it is still possible to aggregate looking times across space and time for analysis, new tools and models allowed researchers to leverage the nature of eye tracking data to answer new questions. Specifically, with these tools, we can now establish where within a stimulus infants look, allowing us to answer questions about how much time infants devote to specific facial features, such as the eyes and mouths of faces (Lewkowicz & Hansen-Tift, 2012; Oakes & Ellis, 2013). We can also present more stimuli simultaneously to infants, better reflecting the attentional demands infants experience as they explore the real world. For instance, we can measure infants’ attention to stimuli, such as faces, in visual search arrays in which multiple stimuli compete for attention at once (e.g., Gliga et al., 2009; Gluckman & Johnson, 2013; Hunter & Markant, 2021; Jakobsen et al., 2016; Kwon et al., 2016). Furthermore, whereas the prior methods emphasized sustained attention mechanisms (i.e., looking time), the temporal precision of eye trackers also yields insight into other aspects of attention, such as orienting (Leppänen, 2016). With measures of where infants looked every 1–33 ms, researchers can determine how quickly infants orient towards specific images. For instance, we can determine when infants shift their gaze to fixate on one of two simultaneously presented stimuli (e.g., Markant & Amso, 2016; Oakes & Kovack-Lesh, 2013; Tummeltshammer et al., 2014), providing insight into processes of attention orienting and attention disengagement. We can measure how quickly they move their eyes to a cued or uncued stimulus in the peripheral visual field (e.g., Ross-Sheehy et al., 2015), providing an understanding of covert attentional processes.
Further advances in eye-tracking, particularly mobile eye-tracking systems, have allowed researchers to examine visual attention in more naturalistic contexts, such as when infants are crawling or walking (Kretch et al., 2014), playing with a caregiver (Kızıldere et al., in press; Yu & Smith, 2013), or manually exploring objects (Chen et al., 2021). Importantly, these systems allow eye tracking from the infants’ own egocentric view, demonstrating how infants deploy attention during learning, in response to the behaviors of others, and as objects come in and out of view. This approach sheds light on both the factors that control infants’ attention in everyday life and the moment-to-moment fluctuations in attention during naturalistic interactions, offering insights that were previously inaccessible with traditional lab-based methods.
2.2. Tools for quantifying stimulus characteristics
Advances in computer vision and modeling during the first 25 years of the 21st century have provided complementary tools to analyze eye-tracking data, allowing for a more nuanced understanding of visual attention development. Specifically, developmental scientists have leveraged advances in computer and cognitive science to quantify image features and relate them to infant attention. This has advanced our knowledge by enabling a more systematic way to test established hypotheses. For example, whereas early work described the effect of physical salience on infant attention using simple visual patterns, such as checkerboards or objects presented in isolation (e.g., Brennan et al., 1966; Cohen et al., 1975; Fantz, 1956, 1963, 1964), we can now present infants with naturalistic scenes and use quantitative models of physical salience (i.e., salience models; Bylinskii et al., 2019; Judd et al., 2012) to more objectively measure low-level stimulus features of the scenes. Some of these salience models were inspired by the properties of V1 and V2 in primate visual cortex (e.g., Harel et al., 2007; Walther & Koch, 2006), allowing us to not only quantify salience in a wide range of displays but also to draw conclusions about the role of early visual cortex in infant looking behavior. The application of these tools has confirmed previous conclusions regarding the decreasing influence of physical salience on infants’ visual attention across the first postnatal year (Frank et al., 2009; Kwon et al., 2016; Pomaranski et al., 2021). Moreover, these tools have led researchers to recognize that the effect of salience is not an all-or-none factor; rather, it is a relative feature (DeBolt et al., 2023), composed of multiple low-level stimulus properties (Hunter et al., 2025), and must be considered within the context of the scene’s semantic features (Amso et al., 2014; Oakes et al., 2024).
The advent of these tools has also paved the way for both researchers of infant and adult attention to examine the influence of a variety of other low- and mid-level stimulus features, such as entropy, feature clutter, texture, and edge density (E. M. Anderson et al., 2024; Balas & Woods, 2014; McAdams et al., 2025; Rosenholtz et al., 2007). As a result, we are no longer limited to only considering whether infants attend to the most salient object in an isolated display. Instead, we also can examine how attention is influenced by a combination of features in more naturalistic contexts. Furthermore, researchers have used these tools to analyze the images taken from cameras mounted on infants’ foreheads (therefore capturing infants’ egocentric fields of view), providing insight into how the properties of scenes that infants actually see vary across development (E. M. Anderson et al., 2024).
New tools have also emerged that allow researchers to quantify higher-order properties of scenes that were once considered too complex to systematically measure. For instance, meaning maps can be created based on adults’ ratings of the semantic informativeness, or meaning, of isolated scene patches (Henderson & Hayes, 2017, 2018). Patches with uniform texture or no identifiable objects are typically rated as low meaning, and those with an identifiable object are often rated as high meaning. These ratings can be combined into a map that reflects the meaningfulness of each location within a naturalistic image. In general, adult fixations are better predicted by these meaning maps than by maps of physical salience (Hayes & Henderson, 2022; Henderson & Hayes, 2017; Peacock et al., 2023). We are just beginning to apply this approach to infants, but the initial findings suggest that infants’ fixations are also influenced by adult meaning ratings and that the effect of meaning increases across the first postnatal year (Oakes et al., 2024).
Finally, researchers have begun to use deep neural network models (DNNs) as tools for modeling how infants’ attention relates to low-, mid-, and high-level feature properties. The models are structured hierarchically, with a sequence of layers that extract increasingly abstract visual information from images, similar to how visual input is processed as it travels from V1 through the inferotemporal cortex in the visual system. In general, DNN layers reasonably approximate activity recorded from the sequence of areas along the adult ventral pathway (e.g., Cadieu et al., 2014; Güçlü & Van Gerven, 2015; Khaligh-Razavi & Kriegeskorte, 2014; Storrs et al., 2020; Wen et al., 2018; Yamins et al., 2014), which are notoriously challenging to measure among awake infant participants (though see Deen et al., 2017). Until higher-resolution fMRI from awake infants is more accessible and feasible, DNNs remain a valuable tool for exploring how different levels of visual abstraction guide attention in early development. For example, Kiat et al. (2022) related the spatial distribution of infants’ fixations to the pattern of activations across units in each layer within a DNN. They found that as infants matured, the spatial distribution of their fixations aligned more closely with the later layers of the DNN, suggesting that their attention was increasingly guided by higher-level visual representations.
3. The emergence of new theories of infant attention development
These innovations in technology and quantitative methods have allowed us to address questions that were unthinkable 25 years ago. These advancements have also reshaped how contemporary researchers studying adults view and discuss attention mechanisms. Specifically, researchers used to argue that adult attention selection is driven by either bottom-up or top-down factors (e.g., Corbetta & Shulman, 2002; Desimone & Duncan, 1995; Egeth & Yantis, 1997; Posner & Petersen, 1990). However, researchers have largely come to the agreement that this dichotomous view is overly simplistic (B. A. Anderson et al., 2021; Awh et al., 2012; Luck et al., 2021) and that these selection mechanisms may be difficult to clearly distinguish (B. A. Anderson, 2024). Recent theories propose that adult visual selective attention is guided by a priority map that integrates multiple factors—such as physical salience, prior learning, motivational relevance, and task goals—to dynamically compute the attentional priority of each location in the visual field (Todd & Manaligod, 2018).
One theoretical advancement in the field of visual attention development is that, similar to adults’, infants’ attention is not controlled by either bottom-up or top-down factors (Markant & Hunter, in press; Markant & Scott, 2018). Although the carefully designed experiments traditionally employed in infant research allowed us to examine bottom-up and top-down selection separately, it has become clear that these mechanisms may be functionally integrated and challenging to disentangle during more complex tasks—such as those afforded by eye-tracking methodologies—or in daily experiences. Thus, the findings of infant attention research over the past 25 years are consistent with the more nuanced theoretical framework adopted by adult attention researchers, indicating that a strict dichotomy between bottom-up versus top-down control cannot fully explain how infants attend. For example, Amso and Scerif (2015) proposed that bottom-up perceptual factors are inextricably linked to higher cognition; attention development occurs as a direct result of infants’ increased ability to process and integrate low-level stimulus information. Markant and Scott (2018) also proposed that biases in face recognition result from a dynamic interaction between bottom-up perceptual and top-down cognitive factors across the first postnatal year. Moreover, empirical work supports these claims: the effect of bottom-up perceptual factors on infants’ attention depends on what they are attending to. Hunter and Markant (2021) reported that when viewing arrays that contained pictures of objects (e.g., shoes, cups, bicycles) and a face, 6- and 11-month-old infants were more likely to look at the most physically salient distractor when the face in the array was an other-race face, compared to when the face in the array was an own-race face. Thus, infants’ attention to physically salient images was influenced by the top-down role of face identity. Similarly, Amso and colleagues (2014) investigated the ability of 4- to 24-month-old infants to detect faces within naturalistic scenes and found that older infants were more likely to detect faces that were more physically salient. In general, such research demonstrates that it may not be possible to understand the isolated roles of bottom-up or top-down factors in infants’ attention, but that researchers should consider how the two factors together determine infants’ attention.
The interaction between top-down and bottom-up attentional control can similarly be seen in how infants’ create their own views of the world through voluntary movements. In one study, 12-month-old infants used controlled head movements to maintain the distance between their head and an attended object, ensuring that the object was larger than surrounding items (Mendez et al., 2024). Thus, infants used top-down attentional control to manipulate the bottom-up physical salience of the objects in their field of view. The conclusion from this work is that the bottom-up versus top-down theoretical framework described by infant visual attention researchers at the beginning of the 21st century may not sufficiently explain the findings generated over the last 25 years. Instead, attention is a complex system, and reflects the interaction of multiple processes (Oakes, 2023b).
A second theoretical advance is the recognition that attention development reflects the cascading effects of multiple dynamically changing systems (Oakes 2023a, Oakes 2023b). Visual attention—just like other abilities such as short-term memory (Ross-Sheehy et al., 2003), motor skills (Malina, 2004), and perception (Kellman & Arterberry, 2000)—does not operate or develop in isolation. Although systems-based approaches to development have long emphasized the interconnectedness of developmental systems (Smith & Thelen, 2003; Thelen & Smith, 1994), such frameworks were not initially integrated into the study of visual attention development. Oakes (2009) argued that attempting to study abilities separately led us to the “Humpty Dumpty problem.” That is, we gained detailed insights into the development of individual processes with little understanding of how they interact in development.
A comprehensive understanding of visual attention development, therefore, requires examining it in the context of other developing systems, as infant vision evolves alongside changes in other domains. Motor milestones, in particular, provide infants with new opportunities to explore and shape their visual experiences. For example, studies equipping infants and toddlers with head-mounted cameras revealed that young infants’ visual experiences are initially dominated by their caregivers’ faces during close contact, such as cradling (Fausey et al., 2016; Jayaraman et al., 2015; Jayaraman & Smith, 2019; Sugden et al., 2014). However, by 9 months, infants spent relatively little time looking at their caregivers’ faces (Abney et al., 2020; Franchak et al., 2011; Suarez-Rivera et al., 2019; Yu & Smith, 2013, 2017). Instead, they focused more on caregivers’ hands and the objects they interacted with. These changes in the content of infants’ visual input reflect the development of postural control and locomotor abilities. Young infants with limited independent movement are typically held by their caregivers and therefore spend much of their time viewing their caregivers’ faces (Fausey et al., 2016). In prone and supine positions, their ability to visually explore their surroundings is restricted by posture and gravity (Adolph & Franchak, 2017). However, as infants develop motor skills, they gain greater control of their posture and head movement, expanding their visual field and creating more variety in their visual experiences (e.g., Kretch et al., 2014; Soska et al., 2010). Kretch and colleagues (2014), for example, demonstrated that infants’ field of view undergoes dramatic changes with the transition from crawling to walking. Importantly, infants’ developing visual attention does not simply reflect changes in motor development. Visual attention also develops alongside changes in infant-caregiver interactions (Karasik et al., 2011, 2014), social cognition (Mundy & Newell, 2007), and object recognition (Reynolds, 2015).
A third theoretical advance is the recognition that visual attention must be understood in context. That is, although significant understanding has been gained from the use of well controlled stimuli with low ecological validity, we may only fully understand the development of visual attention by expanding our understanding to the infants’ everyday lives. Advancements in technology and methods have led researchers to increasingly explore infants’ visual attention to naturalistic stimuli, such as photographs of everyday places and objects (i.e., naturalistic scenes; Oakes et al., 2024; Pomaranski et al., 2021; van Renswoude et al., 2019). Although scene viewing is not precisely the same as attending in a dynamic, 3-dimensional world, the study of infants’ attention to natural scenes has allowed us to make more general conclusions about how infants attend when presented with environments in which many visual stimuli compete for attention. For example, like adults, infants show a strong tendency to fixate on the center of the screen (van Renswoude et al., 2019) and tend to make more saccades across the horizontal axis (van Renswoude et al., 2016). These findings suggest that the biases in fundamental eye movement patterns observed in adults emerge early in development.
Studying attention in context is critical because many studies in the first 25 years of the 21st century demonstrate that visual input shapes the development of visual attention: everyday experiences influence attentional patterns. For example, 4-month-old infants with pets at home focused more on animals’ heads—an informationally rich feature—compared to those without pets (Hurley & Oakes, 2015). The amount of diversity in the faces of infants’ daily experiences influenced how they visually explored own- (familiar) and other- (novel) race faces (Ellis et al., 2017). Furthermore, as infants gain experience, they become increasingly able to control their visual input (Smith et al., 2018). For example, when infants between 18 and 24 months manually played with objects, their self-generated actions created many unique views of a single object which ultimately supported their object recognition and learning of the object names (James, Jones, Smith, et al., 2014; James, Jones, Swain, et al., 2014). Moreover, the limited research that has compared infant visual attention in different cultural contexts suggests that aspects of infants’ everyday lived experience may contribute to how attention develops (DeBolt et al., 2025; Heise et al., 2025).
4. Where do we go next?
The advances we have described thus far lead to the conclusion that understanding the development of visual attention requires more than considering a simple bottom-up and top-down dichotomy and that visual attention must be understood in the context of other developing processes as well as the environment. Infants’ visual attention at any given moment reflects many processes including arousal, perception, stimulus factors, and the infants’ history and goals. Understanding the development of this system requires considering how development in different domains interact. In particular, attention development reflects intertwining cascading changes (Oakes, 2023a), but the exact nature, developmental timescale, and variability across individuals in these changes are still uncharted. Measuring multiple developing systems within and across time in large, diverse samples will therefore allow for a better understanding of how the visual attention cascade unfolds. For instance, researchers may longitudinally measure how low-level stimulus features influence attention before and after infants achieve gross motor skill milestones (e.g., unsupported sitting, crawling, walking), providing insight into how the ability to actively control visual input through motor systems relates to infants’ ability to control their visual attention. Furthermore, infant attention researchers may measure infants’ perceptual capabilities, such as visual acuity or color discrimination thresholds, to directly relate perceptual and cognitive systems (see this approach in children, Lynn et al., 2024).
Furthermore, deeper understanding of infant visual attention will require that researchers measure attention in the real-world, as it dynamically interacts with other systems. This work should complement short-term experimental designs that allow for great control over confounding variables. For example, work by McAdams et al. (2025) found that, when presented with images of building facades, infants preferred looking towards the facades with greater edge orientation entropy. Complementary work by Anderson et al. (2024) indicated that infants’ egocentric views are dominated by simpler edge orientations. Together, this work suggests that McAdams et al.’s (2025) findings may reflect infants’ propensities towards novel stimuli. However, because stimulus features differ in their significance as a function of infants’ past experiences and knowledge, understanding the role of those features on infants’ visual attention requires studying them in context. It is important that future work continues to characterize the properties of infants’ visual inputs to better understand variability across individuals (for examples, see Anderson et al., 2024; Fausey et al., 2016; Kretch et al., 2014).
Finally, it is unclear how much of what we know about infant visual attention is indeed “universal”, and general to infants regardless of culture or type of visual experience, and how much reflects the specific experiences of the populations most commonly studied. One recent study revealed that whereas 6- to 9-month-old infants raised in Northern California were faster to fixate a target, 6- to 9-month-old infants raised in rural Malawi were more accurate in their fixations to the target (DeBolt et al., 2025). Another study found that East Asian 15-month-old infants tended to look longer at items in the background of a scene compared to 15-month-old North American infants (Heise et al., 2025). Differences are also observed in more naturalistic interactions. For example, Chavajay and Rogoff (1999) found that 14- to 20-month-old toddlers from a Guatemalan Mayan community were more likely to attend to multiple events at once, whereas White American toddlers were more likely to shift their attention back and forth between competing events, often focusing on just one event at a time. These studies demonstrate that a more complete understanding of the visual attention development requires including more diverse samples in our research.
To address these limitations in the field, we must increasingly incorporate: 1) longitudinal and multisystem approaches that track infants’ attention over time in the context of other developing skills (e.g., motor skills), social interactions, and environmental changes; 2) ecologically valid stimuli and measures that reflect how infants engage with their environment in natural settings; and finally, 3) samples that represent a broader range of diversity in experience.
Taking this approach will be facilitated by continuing to leverage the advancements in the field of computer science and computer vision to help analyze and interpret the data. We have only just begun to understand how advances in artificial intelligence and machine learning will yield even more promising and sophisticated tools to understand visual attention. For example, using transformers, a deep-learning architecture, DeepMeaning (Hayes & Henderson, in press; Hayes & Henderson, 2023) is a new automated method for identifying objects and semantic content within scenes. One limitation of this approach is that transformer models are trained on image-text relations, and thus the areas of images with the greatest semantic “meaning” identified by the model may not be the most relevant regions from a prelinguistic infant’s perspective. Nevertheless, these tools have the potential to help us move beyond surface-level scene statistics by representing the semantic relations between scene elements (Nuthmann & Henderson, 2010; Peacock et al., 2019).
Similarly, developments in computer vision have resulted in Grounding DINO, a transformer model that facilitates the detection of objects in scenes based on defined object or category labels (Liu et al., 2025). By integrating such tools with eye-tracking or head-mounted cameras, researchers can easily identify which objects infants commonly encounter in everyday interactions, determine the objects they fixate on and how their attention shifts over time, and analyze whether caregivers’ labels (e.g., “ball,” “rattle”) align with infants’ gaze towards objects. Furthermore, Human Pose Estimation (HPE) models have emerged as a computational tool for researchers to automate the analysis of large video datasets. For example, HPE can be used to automatically detect faces, hands, and other features in videos, such as those recorded from head-mounted cameras. Importantly, these models have shown great promise in their ability to replicate and extend well-established findings by characterizing infants’ nonverbal communication (Yurtsever & Eken, 2022) and quantifying how infants’ access to social information dynamically changes across development (Long et al., 2022). Critically, the ease and relative cost of applying these emerging machine-learning tools, compared to traditional behavioral coding methods, make them well-suited for application to datasets from understudied populations, providing greater diversity in our understanding of the relationship between culture and development (Long et al., 2022).
Finally, although infant researchers have begun to relate DNN models to infants’ looking patterns (Kiat et al., 2022), there are several uses of such models that will expand our knowledge. For example, researchers may examine how DNN models relate to different aspects of infants’ visual attention, providing insight into how visual abstraction relates to different components of infant attention (e.g., orienting vs. sustained attention). Furthermore, many existing DNNs do not account for the maturational state of the infant visual system and are not trained in ways that reflect how infants learn in their daily lives (Smith & Slone, 2017). Future models may incorporate biological constraints—such as limited acuity and immature high-level cortex—and be explicitly trained on the types of visual input infants experience. Such models may provide the best insight into the development of visual perception and attention.
5. Conclusion
Clearly, there have been significant shifts in our understanding of the development of visual attention in infancy in the first 25 years of the 21st century. Enabled by advances in technology, modeling, and an understanding of the adult attention system, our conceptualization of how visual attention develops during the first postnatal year has expanded from a relatively simple view to a more nuanced and sophisticated understanding that attention development is an unfolding cascade of events (Oakes, 2023a, 2023b; Oakes & Rakison, 2019). That is, attention development is a dynamic process, shaped by the whole child and their context. Built on the foundation established in the 20th century, we now understand that although the influence of physical salience on attention may decrease over infancy, it may actually operate in the context of priority maps, as has been described for adults (Luck et al., 2021; Todd & Manaligod, 2018). In addition, visual attention cannot be solely understood by examining infants’ looking as they view static, simple stimuli. We now recognize that we must also record infants’ eye gaze as they view complex stimuli (such as photographs and videos of everyday scenes), we must consider the infants’ own egocentric view, and we must study attention development in the context of other developing systems. By investigating attention in these contexts, which previously seemed impossible due to methodological limitations, we have gained important insights into the development of the attentional system, transforming our understanding of how infants perceive and learn from their worlds.
Funding
Preparation of this manuscript was supported by the National Institutes of Health grants 1F32EY034017 (BKH), R01HD108325 (JM), and R01EY030127 (LMO).
Footnotes
CRediT authorship contribution statement
Hunter Brianna: Writing – review & editing, Writing – original draft, Conceptualization. Lisa M. Oakes: Writing – review & editing, Conceptualization. Julie Markant: Writing – review & editing, Conceptualization. Christian M. Nelson: Writing – review & editing, Conceptualization. Shannon M. Klotz: Writing – review & editing. Erim Kızıldere: Writing – review & editing.
Data availability
No data was used for the research described in the article.
References
- Abney DH, Suanda SH, Smith LB, & Yu C (2020). What are the building blocks of parent–infant coordinated attention in free-flowing interaction? Infancy, 25 (6), 871–887. 10.1111/infa.12365 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Adolph KE, & Franchak JM (2017). The development of motor behavior. WIREs Cognitive Science, 8(1–2), Article e1430. 10.1002/wcs.1430 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Amso D, Haas S, & Markant J (2014). An eye tracking investigation of developmental change in Bottom-up attention orienting to faces in cluttered natural scenes. PLoS ONE, 9(1), Article e85701. 10.1371/journal.pone.0085701 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Amso D, & Scerif G (2015). The attentive brain: Insights from developmental cognitive neuroscience. Nature Reviews Neuroscience, 16(10), 606–619. 10.1038/nrn4025 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Amso D, & Tummeltshammer K (2020). Infant visual attention (chapter). In Lockman JJ, & Tamis-LeMonda CS (Eds.), The Cambridge Handbook of Infant Development: Brain, Behavior, and Cultural Context (pp. 186–213). Cambridge: Cambridge University Press. [Google Scholar]
- Anderson BA (2024). Trichotomy revisited: a monolithic theory of attentional control. Vision Research, 217, Article 108366. 10.1016/j.visres.2024.108366 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Anderson BA, Kim H, Kim AJ, Liao M-R, Mrkonja L, Clement A, & Grégoire L (2021). The past, present, and future of selection history. Neuroscience Biobehavioral Reviews, 130, 326–350. 10.1016/j.neubiorev.2021.09.004 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Anderson EM, Candy TR, Gold JM, & Smith LB (2024). An edge-simplicity bias in the visual input to young infants. Science Advances, 10(19), Article eadj8571. 10.1126/sciadv.adj8571 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Aslin RN, & McMurray B (2004). Automated Corneal-Reflection eye tracking in infancy: Methodological developments and applications to cognition. Infancy, 6(2), 155–163. 10.1207/s15327078in0602_1 [DOI] [PubMed] [Google Scholar]
- Atkinson J, & Braddick OJ, & Bremmer G (1989). Development of basic visual functions. In Slater & Bremmer’s edited book “Infant Development”, published by Erlbaum. [Google Scholar]
- Atkinson J, Hood B, Wattam-Bell J, & Braddick O (1996). Changes in infants’ ability to switch visual attention in the first three months of life. Perception, 21(5), 643–653. 10.1068/p210643 [DOI] [PubMed] [Google Scholar]
- Awh E, Belopolsky AV, & Theeuwes J (2012). Top-down versus bottom-up attentional control: A failed theoretical dichotomy. Trends in Cognitive Sciences, 16(8), 437–443. 10.1016/j.tics.2012.06.010 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Balas B, & Woods R (2014). Infant preference for natural texture statistics is modulated by contrast polarity. Infancy: the Official Journal of the International Society on Infant Studies, 19(3), 262–280. 10.1111/infa.12050 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Bornstein MH, Toda S, Azuma H, Tamis-LeMonda C, & Ogino M (1990). Mother and infant activity and interaction in Japan and in the United States: II. A comparative microanalysis of naturalistic exchanges focused on the organisation of infant attention. International Journal of Behavioral Development, 13(3), 289–308. 10.1177/016502549001300303 [DOI] [Google Scholar]
- Bornstein MH, Tamis-LeMonda CS, Pêcheux M-G, & Rahn CW (1991). Mother and infant activity and interaction in France and in the United States: A comparative study. International Journal of Behavioral Development, 14(1), 21–43. 10.1177/016502549101400102 [DOI] [Google Scholar]
- Brennan WM, Ames EW, & Moore RW (1966). Age differences in Infants’ attention to patterns of different complexities. Science, 151(3708), 354–356. 10.1126/science.151.3708.354 [DOI] [PubMed] [Google Scholar]
- Bronson GW (1974). The postnatal growth of visual capacity. Child Development, 45, 873–890. [PubMed] [Google Scholar]
- Bronson GW (1994). Infants’ transitions toward adult-like scanning. Child Development, 65(5), 1243–1261. [DOI] [PubMed] [Google Scholar]
- Bylinskii Z, Judd T, Oliva A, Torralba A, & Durand F (2019). What do different evaluation metrics tell us about saliency models? IEEE Transactions on Pattern Analysis and Machine Intelligence, 41(3), 740–757. 10.1109/TPAMI.2018.2815601 [DOI] [PubMed] [Google Scholar]
- Cadieu CF, Hong H, Yamins DLK, Pinto N, Ardila D, Solomon EA, Majaj NJ, & DiCarlo JJ (2014). Deep neural networks rival the representation of primate IT cortex for core visual object recognition. PLoS Computational Biology, 10(12), Article e1003963. 10.1371/journal.pcbi.1003963 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Chavajay P, & Rogoff B (1999). Cultural variation in management of attention by children and their caregivers. Developmental Psychology, 35(4), 1079–1090. 10.1037//0012-1649.35.4.1079 [DOI] [PubMed] [Google Scholar]
- Chen CH, Houston DM, & Yu C (2021). Parent-Child joint behaviors in novel object play create High-Quality data for word learning. Child Development, 92(5), 1889–1905. 10.1111/cdev.13620 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Clearfield MW, & Jedd KE (2013). The effects of socio-economic status on infant attention. Infant and Child Development, 22(1), 53–67. 10.1002/icd.1770 [DOI] [Google Scholar]
- Cohen LB (1973). A two process model of infant visual attention. Merrill-Palmer Quarterly of Behavior and Development, 19(3), 157–180. [Google Scholar]
- Cohen LB, DeLoache JS, & Rissman MW (1975). The effect of stimulus complexity on infant visual attention and habituation. Child Development, 46(3), 611. 10.2307/1128557 [DOI] [PubMed] [Google Scholar]
- Colombo J (2001). The development of visual attention in infancy. Annual Review of Psychology, 52(1), 337–367. 10.1146/annurev.psych.52.1.337 [DOI] [PubMed] [Google Scholar]
- Colombo J, Richman WA, Shaddy DJ, Follmer Greenhoot A, & Maikranz JM (2001). Heart rate-defined phases of attention, look duration, and infant performance in the paired-comparison paradigm. Child Development, 72(6), 1605–1616. [DOI] [PubMed] [Google Scholar]
- Corbetta M, & Shulman GL (2002). Control of goal-directed and stimulus-driven attention in the brain. Nature Reviews Neuroscience, 3(3), 201–215. 10.1038/nrn755 [DOI] [PubMed] [Google Scholar]
- D’Souza D, Brady D, Haensel JX, & D’Souza H (2020). Is mere exposure enough? The effects of bilingual environments on infant cognitive development. Royal Society Open Science, 7(2), Article 180191. 10.1098/rsos.180191 [DOI] [PMC free article] [PubMed] [Google Scholar]
- DeBolt MC, Mitsven SG, Pomaranski KI, Cantrell LM, Luck SJ, & Oakes LM (2023). A new perspective on the role of physical salience in visual search: Graded effect of salience on infants’ attention. Developmental Psychology, 59(2), 326–343. 10.1037/dev0001460 [DOI] [PMC free article] [PubMed] [Google Scholar]
- DeBolt MC, Caswell BL, George M, Maleta K, Prado EL, Ross-Sheehy S, Stewart CP, & Oakes LM (2025). A Cross-Cultural analysis of Infants’ spatial attention on the infant orienting with attention (IOWA) task. Child Development, 96(3), 1050–1065. 10.1111/cdev.14228 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Deen B, Richardson H, Dilks DD, Takahashi A, Keil B, Wald LL, Kanwisher N, & Saxe R (2017). Organization of high-level visual cortex in human infants. Nature Communications, 8, Article 13995. 10.1038/ncomms13995 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Desimone R, & Duncan J (1995). Neural mechanisms of selective visual attention. Annual Review of Neuroscience, 18(1), 193–222. 10.1146/annurev.ne.18.030195.001205 [DOI] [PubMed] [Google Scholar]
- Egeth HE, & Yantis S (1997). VISUAL ATTENTION: Control, representation, and time course. Annual Review of Psychology, 48(1), 269–297. 10.1146/annurev.psych.48.1.269 [DOI] [PubMed] [Google Scholar]
- Ellis AE, Xiao NG, Lee K, & Oakes LM (2017). Scanning of own- versus other-race faces in infants from racially diverse or homogenous communities. Developmental Psychobiology, 59(5), 613–627. 10.1002/dev.21527 [DOI] [PubMed] [Google Scholar]
- Emberson LL, Cannon G, Palmeri H, Richards JE, & Aslin RN (2017). Using fNIRS to examine occipital and temporal responses to stimulus repetition in young infants: Evidence of selective frontal cortex involvement. Developmental Cognitive Neuroscience, 23, 26–38. 10.1016/j.dcn.2016.11.002 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Fantz RL (1956). A method for studying early visual development. Perceptual and Motor Skills, 6, 13. 10.2466/PMS.6.13-15 [DOI] [Google Scholar]
- Fantz RL (1963). Pattern vision in newborn infants. Science, 140(3564), 296–297. 10.1126/science.140.3564.296 [DOI] [PubMed] [Google Scholar]
- Fantz RL (1964). Visual experience in infants: decreased attention to familiar patterns relative to novel ones. Science, 146(3644), 668–670. 10.1126/science.146.3644.668 [DOI] [PubMed] [Google Scholar]
- Fausey CM, Jayaraman S, & Smith LB (2016). From faces to hands: Changing visual input in the first two years. Cognition, 152, 101–107. 10.1016/j.cognition.2016.03.005 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Franchak JM, Kretch KS, Soska KC, & Adolph KE (2011). Head-Mounted eye tracking: A new method to describe infant looking: Head-Mounted eye tracking. Child Development, 82(6), 1738–1750. 10.1111/j.1467-8624.2011.01670.x [DOI] [PMC free article] [PubMed] [Google Scholar]
- Frank MC, Vul E, & Johnson SP (2009). Development of infants’ attention to faces during the first year. Cognition, 110(2), 160–170. 10.1016/j.cognition.2008.11.010 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Gliga T, Elsabbagh M, Andravizou A, & Johnson M (2009). Faces attract Infants’ attention in complex displays. Infancy, 14(5), 550–562. 10.1080/15250000903144199 [DOI] [PubMed] [Google Scholar]
- Gluckman M, & Johnson SP (2013). Attentional capture by social stimuli in young infants. Frontiers in Psychology, 4. 10.3389/fpsyg.2013.00527 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Greenberg DJ (1977). Visual attention in infancy: processes, methods, and clinical applications. In The structuring of experience (pp. 211–234). Boston, MA: Springer US. [Google Scholar]
- Güçlü U, & Van Gerven MAJ (2015). Deep neural networks reveal a gradient in the complexity of neural representations across the ventral stream. Journal of Neuroscience, 35(27), 10005–10014. 10.1523/JNEUROSCI.5023-14.2015 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Harel J, Koch C, & Perona P (2007). Graph-Based visual saliency. In Schölkopf B, Platt J, & Hofmann T (Eds.), Advances in Neural Information Processing Systems 19 (pp. 545–552). The MIT Press. 10.7551/mitpress/7503.003.0073. [DOI] [Google Scholar]
- Hayes TR, & Henderson JM (2022). Scene inversion reveals distinct patterns of attention to semantically interpreted and uninterpreted features. Cognition, 229, Article 105231. 10.1016/j.cognition.2022.105231 [DOI] [PubMed] [Google Scholar]
- Hayes TR, & Henderson JM (2023). Transformers bridge vision and language to estimate and understand scene meaning. 10.21203/rs.3.rs-2968381/v1. [DOI] [Google Scholar]
- Hayes TR & Henderson JM (in press). DeepMeaning: Estimating and Interpreting Scene Meaning for Attention Using a Vision-Language Transformer. Open Mind: Discoveries in Cognitive Science. [Google Scholar]
- Heise MJ, Meristo M, Ueno M, Itakura S, & Carlson SM (2025). Cultural differences in visual attention emerge in infancy. Infancy, 30(1), Article e12651. 10.1111/infa.12651 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Henderson JM, & Hayes TR (2017). Meaning-based guidance of attention in scenes as revealed by meaning maps. Nature Human Behaviour, 1(10), 743–747. 10.1038/s41562-017-0208-0 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Henderson JM, & Hayes TR (2018). Meaning guides attention in real-world scene images: Evidence from eye movements and meaning maps. Journal of Vision, 18 (6), 10. 10.1167/18.6.10 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Hunter BK, Kiat JE, Klotz SM, Nelson CM, Luck SJ, & Oakes LM (2025). The predictive ability of GBVS feature channels on infants’ fixations of natural scenes. Visual Cognition, 1–19. 10.1080/13506285.2025.2468690 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Hunter BK, & Markant J (2021). Differential sensitivity to species- and race-based information in the development of attention orienting and attention holding face biases in infancy. Developmental Psychobiology, 63(3), 461–469. 10.1002/dev.22027 [DOI] [PubMed] [Google Scholar]
- Hurley KB, & Oakes LM (2015). Experience and distribution of attention: Pet exposure and Infants’ scanning of animal images. Journal of Cognition and Development, 16(1), 11–30. 10.1080/15248372.2013.833922 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Jakobsen KV, Umstead L, & Simpson EA (2016). Efficient human face detection in infancy. Developmental Psychobiology, 58(1), 129–136. 10.1002/dev.21338 [DOI] [PubMed] [Google Scholar]
- James KH, Jones SS, Smith LB, & Swain SN (2014). Young Children’s Self-Generated object views and object recognition. Journal of Cognition and Development, 15(3), 393–401. 10.1080/15248372.2012.749481 [DOI] [PMC free article] [PubMed] [Google Scholar]
- James KH, Jones SS, Swain S, Pereira A, & Smith LB (2014). Some views are better than others: evidence for a visual bias in object views self-generated by toddlers. Developmental Science, 17(3), 338–351. 10.1111/desc.12124 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Jayaraman S, Fausey CM, & Smith LB (2015). The faces in Infant-Perspective scenes change over the first year of life. PLOS ONE, 10(5), Article e0123780. 10.1371/journal.pone.0123780 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Jayaraman S, & Smith LB (2019). Faces in early visual environments are persistent not just frequent. Vision Research, 157, 213–221. 10.1016/j.visres.2018.05.005 [DOI] [PubMed] [Google Scholar]
- Johnson MH (1990). Cortical maturation and the development of visual attention in early infancy. Journal of Cognitive Neuroscience, 2(2), 81–95. 10.1162/jocn.1990.2.2.81 [DOI] [PubMed] [Google Scholar]
- Johnson MH (1995). The inhibition of automatic saccades in early infancy. Developmental Psychobiology: The Journal of the International Society for Developmental Psychobiology, 28(5), 281–291. [DOI] [PubMed] [Google Scholar]
- Judd T, Durand F, & Torralba A (2012). A Benchmark of Computational Models of Saliency to Predict Human Fixations. 〈https://dspace.mit.edu/handle/1721.1/68590〉.
- Karasik LB, Tamis-LeMonda CS, & Adolph KE (2011). Transition from crawling to walking and Infants’ actions with objects and people: Crawling, walking, and objects. Child Development, 82(4), 1199–1209. 10.1111/j.1467-8624.2011.01595.x [DOI] [PMC free article] [PubMed] [Google Scholar]
- Karasik LB, Tamis-LeMonda CS, & Adolph KE (2014). Crawling and walking infants elicit different verbal responses from mothers. Developmental Science, 17(3), 388–395. 10.1111/desc.12129 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Kellman PJ, & Arterberry ME (2000). The cradle of knowledge: Development of perception in infancy. MIT press. [Google Scholar]
- Khaligh-Razavi S-M, & Kriegeskorte N (2014). Deep supervised, but not unsupervised, models May explain IT cortical representation. PLoS Computational Biology, 10(11), Article e1003915. 10.1371/journal.pcbi.1003915 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Kızıldere E, & Nelson CM, Casasola M, Graf Estes K, Oakes LM (in press). Parents’ multimodal spatial language structures infants’ in-the-moment attention during spatial play. Developmental Psychology. [DOI] [PubMed] [Google Scholar]
- Kiat JE, Luck SJ, Beckner AG, Hayes TR, Pomaranski KI, Henderson JM, & Oakes LM (2022). Linking patterns of infant eye movements to a neural network model of the ventral stream using representational similarity analysis. Developmental Science, 25(1), Article e13155. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Kretch KS, Franchak JM, & Adolph KE (2014). Crawling and walking infants see the world differently. Child Development, 85(4), 1503–1518. 10.1111/cdev.12206 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Kwon M-K, Setoodehnia M, Baek J, Luck SJ, & Oakes LM (2016). The development of visual search in infancy: Attention to faces versus salience. Developmental Psychology, 52(4), 537–555. 10.1037/dev0000080 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Leppänen JM (2016). Using eye tracking to understand Infants’ attentional bias for faces. Child Development Perspectives, 10(3), 161–165. 10.1111/cdep.12180 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Lewkowicz DJ, & Hansen-Tift AM (2012). Infants deploy selective attention to the mouth of a talking face when learning speech. Proceedings of the National Academy of Sciences, 109(5), 1431–1436. 10.1073/pnas.1114783109 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Liu S, Zeng Z, Ren T, Li F, Zhang H, Yang J, Jiang Q, Li C, Yang J, Su H, Zhu J, & Zhang L (2025). Grounding DINO: marrying DINO with grounded Pre-training for Open-Set object detection. In Leonardis A, Ricci E, Roth S, Russakovsky O, Sattler T, & Varol G (Eds.), Computer Vision – ECCV 2024, 15105 pp. 38–55). Springer Nature Switzerland. 10.1007/978-3-031-72970-6_3. [DOI] [Google Scholar]
- Long BL, Kachergis G, Agrawal K, & Frank MC (2022). A longitudinal analysis of the social information in infants’ naturalistic visual experience using automated detections. Developmental Psychology, 58(12), 2211–2229. 10.1037/dev0001414 [DOI] [PubMed] [Google Scholar]
- Luck SJ, Gaspelin N, Folk CL, Remington RW, & Theeuwes J (2021). Progress toward resolving the attentional capture debate. Visual Cognition, 29(1), 1–21. 10.1080/13506285.2020.1848949 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Lynn A, Maule J, & Amso D (2024). Visual and cognitive processes contribute to age-related improvements in visual selective attention. Child Development, 95(2), 391–408. 10.1111/cdev.13992 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Maia MG, Soker-Elimaliah S, Jancart K, Harbourne RT, & Berger SE (2024). Focused attention as a new sitter: How do infants balance it all? Infant Behavior Development, 74, Article 101926. 10.1016/j.infbeh.2024.101926 [DOI] [PubMed] [Google Scholar]
- Malina RM (2004). Motor development during infancy and early childhood: Overview and suggested directions for research. International Journal of Sport and Health Science, 2, 50–66. [Google Scholar]
- Markant J, & Amso D (2016). The development of selective attention orienting is an agent of change in learning and memory efficacy. Infancy: the Official Journal of the International Society on Infant Studies, 21(2), 154–176. 10.1111/infa.12100 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Markant J, & Hunter BK (in press). Visual selective attention. Oxford Handbook of Perceptual Development. [Google Scholar]
- Markant J, & Scott LS (2018). Attention and perceptual learning interact in the development of the Other-Race effect. Current Directions in Psychological Science, 27 (3), 163–169. 10.1177/0963721418769884 [DOI] [Google Scholar]
- McAdams P, Svobodova S, Newman TJ, Terry K, Mather G, Skelton AE, & Franklin A (2025). The edge orientation entropy of natural scenes is associated with infant visual preferences and adult aesthetic judgements. PloS One, 20(2), Article e0316555. 10.1371/journal.pone.0316555 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Mendez AH, Yu C, & Smith LB (2024). Controlling the input: How one-year-old infants sustain visual attention. Developmental Science, 27(2), Article e13445. 10.1111/desc.13445 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Mundy P, & Newell L (2007). Attention, joint attention, and social cognition. Current Directions in Psychological Science, 16(5), 269–274. 10.1111/j.1467-8721.2007.00518.x [DOI] [PMC free article] [PubMed] [Google Scholar]
- Nuthmann A, & Henderson JM (2010). Object-based attentional selection in scene viewing. –20 Journal of Vision, 10(8), 20. 10.1167/10.8.20. [DOI] [PubMed] [Google Scholar]
- Oakes LM (2009). The “Humpty dumpty Problem” in the study of early cognitive development: Putting the infant back together again. Perspectives on Psychological Science, 4(4), 352–358. 10.1111/j.1745-6924.2009.01137.x [DOI] [PMC free article] [PubMed] [Google Scholar]
- Oakes LM (2012). Advances in eye tracking in infancy research. Infancy, 17(1), 1–8. 10.1111/j.1532-7078.2011.00101.x [DOI] [PubMed] [Google Scholar]
- Oakes LM (2023a). The cascading development of visual attention in infancy: Learning to look and looking to learn. Current Directions in Psychological Science, 32(5), 410–417. 10.1177/09637214231178744 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Oakes LM (2023b). The development of visual attention in infancy: A cascade approach. In In Advances in Child Development and Behavior, 64 pp. 1–37). Elsevier. 10.1016/bs.acdb.2022.10.004 [DOI] [PubMed] [Google Scholar]
- Oakes LM, & Ellis AE (2013). An Eye-Tracking investigation of developmental changes in Infants’ exploration of upright and inverted human faces. Infancy, 18 (1), 134–148. 10.1111/j.1532-7078.2011.00107.x [DOI] [PMC free article] [PubMed] [Google Scholar]
- Oakes LM, Hayes TR, Klotz SM, Pomaranski KI, & Henderson JM (2024). The role of local meaning in infants’ fixations of natural scenes. Infancy, 29(2), 284–298. 10.1111/infa.12582 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Oakes LM, & Kovack-Lesh KA (2013). Infants’ visual recognition memory for a series of categorically related items. Journal of Cognition and Development, 14(1), 63–86. 10.1080/15248372.2011.645971 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Oakes LM, & Rakison DH (2019). Developmental cascades: Building the infant mind. Oxford University Press. [Google Scholar]
- Peacock CE, Hayes TR, & Henderson JM (2019). Meaning guides attention during scene viewing, even when it is irrelevant. Attention, Perception Psychophysics, 81(1), 20–34. 10.3758/s13414-018-1607-7 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Peacock CE, Singh P, Hayes TR, Rehrig G, & Henderson JM (2023). Searching for meaning: Local scene semantics guide attention during natural visual search in scenes. Quarterly Journal of Experimental Psychology, 76(3), 632–648. 10.1177/17470218221101334 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Pickron CB, Iyer A, Fava E, & Scott LS (2018). Learning to individuate: The specificity of labels differentially impacts infant visual attention. Child Development, 89(3), 698–710. 10.1111/cdev.13004 [DOI] [PubMed] [Google Scholar]
- Pomaranski KI, Hayes TR, Kwon M-K, Henderson JM, & Oakes LM (2021). Developmental changes in natural scene viewing in infancy. Developmental Psychology, 57(7), 1025–1041. 10.1037/dev0001020 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Posner MI (1980). Orienting of attention. Quarterly Journal of Experimental Psychology, 32(1), 3–25. 10.1080/00335558008248231 [DOI] [PubMed] [Google Scholar]
- Posner MI, & Petersen SE (1990). The attention system of the human brain. Annual Review of Neuroscience, 13(1), 25–42. 10.1146/annurev.ne.13.030190.000325 [DOI] [PubMed] [Google Scholar]
- Reynolds GD (2015). Infant visual attention and object recognition. Behavioural brain Research, 285, 34–43. 10.1016/j.bbr.2015.01.015 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Reynolds GD, & Richards JE (2005). Familiarization, attention, and recognition memory in infancy: An event-related potential and cortical source localization study. Developmental Psychology, 41(4), 598. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Reynolds GD, & Romano AC (2016). The development of attention systems and working memory in infancy. Frontiers in Systems Neuroscience, 10(15). 10.3389/fnsys.2016.00015 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Rogoff B (1990). Apprenticeship in thinking: cognitive development in social context. New York: Oxford University Press. [Google Scholar]
- Rosenholtz R, Li Y, & Nakano L (2007). Measuring visual clutter. Journal of Vision, 7(2), 17. 10.1167/7.2.17 [DOI] [PubMed] [Google Scholar]
- Ross-Sheehy S, Oakes LM, & Luck SJ (2003). The development of visual short-term memory capacity in infants. Child Development, 74(6), 1807–1822. [DOI] [PubMed] [Google Scholar]
- Ross-Sheehy S, Schneegans S, & Spencer JP (2015). The infant orienting with attention task: Assessing the neural basis of spatial attention in infancy. Infancy, 20 (5), 467–506. 10.1111/infa.12087 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Ruff HA, & Turkewitz G (1975). Developmental changes in the effectiveness of stimulus intensity on infant visual attention. Developmental Psychology, 11(6), 705. [Google Scholar]
- Singh L, Fu CS, Rahman AA, Hameed WB, Sanmugam S, Agarwal P, Jiang B, Chong YS, Meaney MJ, Rifkin-Graboi A, & GUSTO Research Team. (2015). Back to basics: A bilingual advantage in infant visual habituation. Child Development, 86(1), 294–302. 10.1111/cdev.12271 [DOI] [PubMed] [Google Scholar]
- Smith LB, & Gasser M (2005). The development of embodied cognition: Six lessons from babies. Artificial Life, 11(1–2), 13–29. 10.1162/1064546053278973 [DOI] [PubMed] [Google Scholar]
- Smith LB, Jayaraman S, Clerkin E, & Yu C (2018). The developing infant creates a curriculum for statistical learning. Trends in Cognitive Sciences, 22(4), 325–336. 10.1016/j.tics.2018.02.004 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Smith LB, & Slone LK (2017). A developmental approach to machine learning? Frontiers in Psychology, 8, 2124. 10.3389/fpsyg.2017.02124 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Smith LB, & Thelen E (2003). Development as a dynamic system. Trends in Cognitive Sciences, 7(8), 343–348. 10.1016/S1364-6613(03)00156-6 [DOI] [PubMed] [Google Scholar]
- Soska KC, Adolph KE, & Johnson SP (2010). Systems in development: Motor skill acquisition facilitates three-dimensional object completion. Developmental Psychology, 46(1), 129–138. 10.1037/a0014618 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Storrs KR, Khaligh-Razavi S-M, & Kriegeskorte N (2020). Noise ceiling on the crossvalidated performance of reweighted models of representational dissimilarity: Addendum to Khaligh-Razavi & Kriegeskorte (2014). 10.1101/2020.03.23.003046. [DOI] [Google Scholar]
- Suarez-Rivera C, Smith LB, & Yu C (2019). Multimodal parent behaviors within joint attention support sustained attention in infants. Developmental Psychology, 55(1), 96–109. 10.1037/dev0000628 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Sugden NA, Mohamed-Ali MI, & Moulson MC (2014). I spy with my little eye: Typical, daily exposure to faces documented from a first-person infant perspective. Developmental Psychobiology, 56(2), 249–261. 10.1002/dev.21183 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Thelen E, & Smith LB (1994). A dynamic systems approach to the development of cognition and action. The MIT Press. [Google Scholar]
- Todd RM, & Manaligod MGM (2018). Implicit guidance of attention: The priority state space framework. Cortex, 102, 121–138. 10.1016/j.cortex.2017.08.001 [DOI] [PubMed] [Google Scholar]
- Treisman A (1964). Selective attention in man. British Medical Bulletin, 20(1), 12–16. 10.1093/oxfordjournals.bmb.a070274 [DOI] [PubMed] [Google Scholar]
- Tummeltshammer KS, Mareschal D, & Kirkham NZ (2014). Infants’ selective attention to reliable visual cues in the presence of salient distractors. Child Development, 85(5), 1981–1994. 10.1111/cdev.12239 [DOI] [PubMed] [Google Scholar]
- van Renswoude DR, Johnson SP, Raijmakers MEJ, & Visser I (2016). Do infants have the horizontal bias? Infant Behavior and Development, 44, 38–48. 10.1016/j.infbeh.2016.05.005 [DOI] [PubMed] [Google Scholar]
- van Renswoude DR, van den Berg L, Raijmakers MEJ, & Visser I (2019). Infants’ center bias in free viewing of real-world scenes. Vision Research, 154, 44–53. 10.1016/j.visres.2018.10.003 [DOI] [PubMed] [Google Scholar]
- Vygotsky L (1978). Mind in society: The development of higher psychological processes. Cambridge, MA: Harvard University Press. [Google Scholar]
- Walther D, & Koch C (2006). Modeling attention to salient proto-objects. Neural Networks, 19(9), 1395–1407. 10.1016/j.neunet.2006.10.001 [DOI] [PubMed] [Google Scholar]
- Wen H, Shi J, Chen W, & Liu Z (2018). Deep residual network predicts cortical representation and organization of visual features for rapid categorization. Scientific Reports, 8(1), 3752. 10.1038/s41598-018-22160-9 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Xie W, & Richards JE (2016). Effects of interstimulus intervals on behavioral, heart rate, and event-related potential indices of infant engagement and sustained attention. Psychophysiology, 53(8), 1128–1142. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Yamins DLK, Hong H, Cadieu CF, Solomon EA, Seibert D, & DiCarlo JJ (2014). Performance-optimized hierarchical models predict neural responses in higher visual cortex. Proceedings of the National Academy of Sciences, 111(23), 8619–8624. 10.1073/pnas.1403112111 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Yu C, & Smith LB (2013). Joint attention without gaze following: human infants and their parents coordinate visual attention to objects through Eye-Hand coordination. PLoS ONE, 8(11), Article e79659. 10.1371/journal.pone.0079659 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Yu C, & Smith LB (2017). Hand–Eye coordination predicts joint attention. Child Development, 88(6), 2060–2078. 10.1111/cdev.12730 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Yurtsever MME, & Eken S (2022). BabyPose: Real-Time decoding of Baby’s Non-Verbal communication using 2D Video-Based pose estimation. IEEE Sensors Journal, 22(14), 13776–13784. 10.1109/JSEN.2022.3183502 [DOI] [Google Scholar]
- Zhang Y, Martinez-Cedillo AP, Mason HT, Vuong QC, Garcia-de-Soria MC, Mullineaux D, Knight MI, & Geangu E (2025). An automatic sustained attention prediction (ASAP) method for infants and toddlers using wearable device signals. Scientific Reports, 15(1), Article 13298. 10.1038/s41598-025-96794-x [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Data Availability Statement
No data was used for the research described in the article.
