Abstract
Children tend to explore broadly, and despite its costs, exploration confers important benefits, especially early in development when people know little. What drives this exploratory behavior? In contrast to accounts emphasizing heightened curiosity or increased randomness, we propose that early exploration reflects the immaturity of the working memory-selective attention system.
Keywords: Cognitive Development, Exploration, Working Memory, Attention
Exploration, exploitation, and development
Human intelligence is characterized by a remarkable ability to learn, allowing individuals to adapt to a wide range of environments. However, learners face an important dilemma: repeat actions proven to be successful or try new actions? The choice is not obvious because neither option is cost-free: trying something new may waste time and effort, whereas repeating familiar actions may forgo the opportunity to expand knowledge. This problem is known as the exploration-exploitation dilemma [1], which has no general optimal solution. Unlike adults who are typically more exploitative, children tend to explore broadly [2,3], and exploration has been shown to be beneficial, especially early in development. What drives early exploratory behavior? We argue that early exploration is a direct consequence of the immaturity of the working memory-selective attention (WM-SA) system.
Benefits of early exploration
To appreciate the benefits of early exploration, consider the alternative: relatively narrow information sampling driven by existing knowledge, which is a hallmark of mature cognition. Such an exploitative pattern of sampling often results in learning traps – missing information that could prove useful. Imagine discovering, after a considerable search, gas stations with lower gas prices. Although sampling new stations could be disappointing, staying strictly with the known ones may result in failing to discover a new, cheaper gas station. In fact, “recommender algorithms” deployed by social media companies create such learning traps by steering users toward content for which they have shown a preference [4]. By contrast, broad information sampling is less efficient, but it allows learners to build a more complete representation of the environment, which is critical early in development. Thus, the primary benefit of exploration is avoiding learning traps by learning more than is learned under narrow sampling. Consequently, under many circumstances, children notice what adults miss [5–7], which happens when adults try to both optimize performance and reduce sampling and decision complexity by deploying selective attention. Importantly, this tendency to reduce complexity while optimizing performance is likely to stem from the accuracy and efficiency goals: When these are removed in a goal-free task, thus increasing uncertainty about what information to sample, older child participants tend to explore more broadly than younger ones [8].
What drives early exploratory behavior?
Recognizing the need for exploration in learning, computer scientists have developed several ways of facilitating exploratory actions, including “exploration bonuses” – rewarding occasional exploratory actions [9] and the epsilon-greedy algorithm, allowing random exploratory actions. Although these ideas have corresponding analogs in theories of human exploratory behavior (e.g., uncertainty-based and random exploration, respectively), there is also a profound difference: whereas the computer science literature focuses primarily on how to implement exploratory behavior in artificial systems, the human learning literature focuses primarily on what drives this behavior in natural systems.
Researchers have generally discussed three potential mechanisms driving exploratory behavior early in human development [3,10]. First, young children may be highly curious, and curiosity drives exploration. Second, immature systems may tend to accumulate perceptual, processing, and decisional noise, leading to less deterministic, more random, less exploitative decisions. And third, exploration in children may arise from the immaturity of the WM-SA system.
Although not necessarily mutually exclusive, these three accounts make distinct predictions about sampling patterns in learning and decision-making tasks: whereas under the curiosity-driven exploration account broad sampling reflects the tendency to maximize information gain or learning, under both the randomness and WM-SA immaturity accounts, broad sampling reflects an inability to sample information efficiently. However, whereas randomness involves sampling with no temporal structure, WM-SA immaturity involves sampling with distinct temporal structure [10]. Specifically, an immature WM-SA system is prone to some time-dependent forgetting of sampling decisions and outcomes, resulting in increasing epistemic uncertainty about or novelty of options that have not been resampled recently. This uncertainty or novelty, in turn, biases sampling toward options with longer sampling lags. Recent research supports WM-SA immaturity as a likely mechanism of early exploration, while undermining other candidate accounts [3,10].
In our theoretical work [11,12], we consider how the mature WM-SA system may support increasingly efficient performance and the narrowing of attentional focus resulting in exploitative behavior (Figure 1A). The underlying idea is that even prior to embarking on a task, the learner has a WM representation of their goals and features of the task that may be particularly useful (i.e., a “feature priority map”). This representation directs attention toward the goals and features of greater perceived priority or relevance, with other features potentially inhibited. Once an action is performed and the outcome (which may confirm or violate the learner’s expectations) is observed, the WM representation is strengthened or revised, and the cycle repeats. These processing cycles result in progressively more efficient choices and progressively narrower attentional focus, both manifesting as exploitative behavior.
Figure 1. Operation of the WM-SA system.

LTM (long-term memory) – general semantic knowledge about situations, environments, and tasks. Based on this general knowledge, people form situation-specific expectations. These general expectations, along with task goals and maps are represented in WM. WM guides attention toward more important and more relevant aspects of the task and the stimuli, making attention selective. Selective attention allows participants to (a) focus on the relevant aspects of the task and (b) inhibit the irrelevant ones. Post-decision attention focuses on the difference between an expected and observed outcome of a trial. It is critical for WM updating and subsequent strengthening or revision of WM representation. A. Mature System. B. Immature system. Dashed lines represent potential sources of immaturity (WM guidance, Inhibition of distracters, Error detection, and WM updating).
However, the case is different when the WM-SA system is immature (see Figure 1B, with dashed lines identifying potential sources of immaturity). There are several potential, non-mutually exclusive sources of immaturity. First, WM representations may be too imprecise or weak to effectively guide SA. Second, the SA system may struggle to inhibit actions or features that have led to poorer outcomes. Third, WM may fail to update based on the outcomes of previous actions.
The first possibility has been supported by recent research focusing on patterns of choices in a 4-armed bandit decision-making task and on attention allocation in a category learning task [10,13]. To directly test the role of WM guidance of SA in information sampling, we taxed adults’ WM resources using a dual-task paradigm (a concurrent verbal WM load). While not a direct model of development, this manipulation provides a functional analogy to the immature WM-SA system, allowing us to test whether reducing WM resources produces broader information sampling and exploratory behavior, like those observed in children. Notably, WM load (unless it overwhelms the learner) may specifically affect attentional distribution rather than impairing learning and performance, more generally.
The overall task is presented in Figure 2A. Adults performed a simple decision-making task with or without WM load. Additionally, a group of 5-year-olds without WM load was included as a comparison group. Exploitation was inferred from participants’ predominantly selecting the highest-reward option, whereas exploration was inferred from a reduced dominance of the high-reward option and participants’ tendency to switch among options.
Figure 2. Effects of WM load on exploratory behavior.

Adults under WM load (Dual-task condition), without WM load (Control conditions), and 5-year-olds without WM load performed a 4-armed bandit decision making tasks. Proportions of choices of different-valued options, proportions of trial-to-trial switching, and best-fitting models of trial-trial-behavior are compared across these groups. Similar to 5-year-old children, dual-task adults’ performance is more exploratory than that of the control adults. A. Decision-making task under the Control and Dual Task conditions (after [10]). Each creature represents a different reward value (1, 2, 3, and 10 candies), and the participants’ task is to collect as many candies as possible. B. Proportion of choice options (values) across conditions. C. Response switching rates across conditions. D. Model-fitting results: Best-fitting models by number of participants and conditions.
Choice proportions are shown in Figure 2B, and switching rates are shown in Figure 2C. Critically, whereas adults in the control condition exhibited exploitative behavior, adults in the WM load condition, like children, exhibited more exploratory behavior. We also fit individuals’ patterns of behavior using computational models implementing value-based exploitation, uncertainty-based exploration (i.e., choosing options that had gone unsampled the longest), and random responding. Whereas all adults in the control condition were best fit by the exploitation model, a large proportion of young children and dual-task adults were best fit by the uncertainty-driven exploration model (see Figure 2D). Therefore, taxing adults’ WM made their behavior more exploratory, yet not random [10]. Similar patterns were found in a category learning task under WM load: adults under WM load sampled a broader range of stimulus dimensions, mirroring the distributed attention patterns previously documented in young children. These results implicate WM guidance in exploitative behavior, narrow information sampling, and attentional selectivity, and are consistent with the hypothesis that exploratory behaviors stem from the immaturity of cognitive control. Future research will examine the contribution of other components of the WM-SA system to exploration-exploitation behavior.
Concluding remarks
Recent work suggests that exploratory behavior may arise in part from the immaturity of the WM-SA system. We believe that such broad information sampling, distributed attention, and other variants of exploratory behavior have important benefits, especially early in development, when learners know little about their environments. Although a typical treatment of the exploration-exploitation dilemma in artificial systems is through randomness or “exploration bonuses”[9], our work suggests that natural systems implement a different solution – protracted immaturity resulting in weakened WM guidance of selective attention.
Acknowledgments
We thank John Opfer, Layla Unger, and Sami Yousif for providing helpful feedback on this manuscript. Research reported here and writing of this manuscript have been supported by National Institutes of Health Grants R01HD078545 and R01HD111458 to Vladimir M. Sloutsky.
Glossary
- Distributed attention
Distributed attention is the broad allocation of attention across multiple features, locations, stimulus dimensions or objects, resulting in less selective information sampling
- Exploitation
Repeating actions that have previously led to a successful, rewarding outcome
- Exploration
Performing new actions, whose outcomes are unknown or not remembered
- Exploration/Exploration Dilemma
Reflects the fact that neither option is cost-free. Exclusive exploitation carries a risk of missing important knowledge, whereas exclusive exploration carries a risk of missing material rewards
- Selective Attention
Selective attention is the preferential processing of behaviorally relevant information at the expense of less relevant information
- Working Memory
Working memory is a limited-capacity system that actively maintains task-relevant information to guide attention, decision making, and action
Footnotes
Publisher's Disclaimer: This is a PDF file of an unedited manuscript that has been accepted for publication. As a service to our customers we are providing this early version of the manuscript. The manuscript will undergo copyediting, typesetting, and review of the resulting proof before it is published in its final form. Please note that during the production process errors may be discovered which could affect the content, and all legal disclaimers that apply to the journal pertain.
References
- 1.Hills TT, Todd PM, Lazer D, Redish AD, Couzin ID, Group the CSR. Exploration versus exploitation in space, mind, and society. Trends Cogn Sci. 2015;19(1):46–54. doi: 10.1016/j.tics.2014.10.004. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 2.Gopnik A. Childhood as a solution to explore-exploit tensions. Philosophical Transactions of the Royal Society B: Biological Sciences. 2020. Jul 20;375(1803):20190502. doi: 10.1098/rstb.2019.0502. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3.Blanco NJ, Sloutsky VM. Exploration, exploitation, and development: Developmental shifts in decision-making. Child Dev. 2024. doi: 10.1111/cdev.14070. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4.Bahg G, Sloutsky VM, Turner BM. Algorithmic personalization of information can cause inaccurate generalization and overconfidence. J Exp Psychol: Gen. 2025;154(9):2503–22. doi: 10.1037/xge0001763. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 5.Blanco NJ, Turner BM, Sloutsky VM. The benefits of immature cognitive control: How distributed attention guards against learning traps. J Exp Child Psychol. 2023;226:105548. doi: 10.1016/j.jecp.2022.105548. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6.Plebanek DJ, Sloutsky VM. Costs of selective attention: when children notice what adults miss. Psych Sci. 2017;28(6):723–32. doi: 10.1177/0956797617693005 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7.Liquin EG, Gopnik A. Children are more exploratory and learn more than adults in an approach-avoid task. Cognition. 2022;218:104940. doi: 10.1016/j.cognition.2021.104940. [DOI] [PubMed] [Google Scholar]
- 8.Pelz M, Kidd C. The elaboration of exploratory play. Philosophical Transactions of the Royal Society B: Biological Sciences. 2020. Jul 20;375(1803):20190503. doi: 10.1098/rstb.2019.0503. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9.Dayan P, Sejnowski TJ. Exploration bonuses and dual control. Mach Learn. 1996;25(1):5–22. doi: 10.1023/a:1018357105171 [DOI] [Google Scholar]
- 10.Wan Q, Sloutsky VM. Working memory shapes information sampling and attention allocation across development. J Exp Psychol: Gen. 2026;155(2):479–98. doi: 10.1037/xge0001848. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11.Turner BM, Sloutsky VM. Cognitive Inertia: Cyclical Interactions Between Attention and Memory Shape Learning. Curr Dir Psychol Sci. 2024;33(2):79–86. doi: 10.1177/09637214231217989 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12.Galdo M, Weichart ER, Sloutsky VM, Turner BM. The quest for simplicity in human learning: Identifying the constraints on attention. Cogn Psychol. 2022;138:101508. doi: 10.1016/j.cogpsych.2022.101508. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13.Wan Q, Sloutsky VM. Exploration, distributed attention, and development of category learning. Psychol Sci. 2024;35(10):1164–77. doi: 10.1177/09567976241258146. [DOI] [PMC free article] [PubMed] [Google Scholar]
