Skip to main content
Proceedings of the National Academy of Sciences of the United States of America logoLink to Proceedings of the National Academy of Sciences of the United States of America
. 2025 Oct 27;122(44):e2416720122. doi: 10.1073/pnas.2416720122

Learning expectations shape cognitive control allocation

Javier Alejandro Masís Obando a,1, Sebastian Musslick b,c, Jonathan D Cohen a,d
PMCID: PMC12595486  PMID: 41144677

Significance

People learn constantly, over timescales short (memorizing game rules) and long (going to law school). As commonplace as learning is, however, investing time to learn instead of engaging in already familiar activities often violates two core tenets of current theories of mental effort: choosing to do something that is both less immediately rewarding and cognitively more effortful. Violating these tenets presumably reflects the recognition that learning will be beneficial later. We formalize this hypothesis in an augmented theory of mental effort that takes into account the future value of learning, and show—in the first controlled study of mental effort allocation in the service of learning—that people are willing to allocate more mental effort when they expect to learn.

Keywords: learning, decision making, cognitive control, expected value of control, drift diffusion model

Abstract

Current models frame the allocation of cognitive control as a process of expected utility maximization. The benefits of a candidate control signal are weighed against its costs (e.g., opportunity costs). Recent theorizing has found that, despite promoting the counterintuitive behavior of longer deliberation, which is less rewarding in the short term, it is nevertheless normative to account for the value of learning when determining control allocation. Here, we sought to test this proposal by examining whether people were willing to allocate greater control and thereby expend greater effort (e.g., deliberate for longer) when they perceived a task to be learnable compared to when they did not. We found that participants’ proficiency and learning rate in the first block of a simple perceptual dot-motion task were able to predict their willingness to deliberate in a second block. These findings support the hypothesis that agents consider learnability when allocating cognitive control, and comply with a formal model of control allocation that considers the future discounted value of learning on reward.


Anyone who regularly types on a keyboard knows that typing technique falls into two broad categories: easy (hunting and pecking) and hard (touch typing). Why would anyone pick the hard way? Because, with enough practice, the hard way will lead to faster typing (1), a better result in the long term. In fact, with enough practice, the relatively hard technique of touch typing will cease to be hard as it transitions from a control-dependent process to an automatic one (24).

At the heart of this decision, and countless others like it, is a form of intertemporal choice that we face throughout our lives on timescales big and small. For example, when deciding whether to choose the more immediately useful but less efficient strategy (hunting and pecking) or to invest the time and mental effort in learning the less immediately useful but ultimately more efficient one (touch typing), a novice should consider how long into the future they will be performing (typing), how quickly they can gain proficiency, and how important it is that they are proficient (e.g., whether they will be paid for it).

These choices range from the quotidian (make pasta again or follow a new recipe) to the life-changing (get a job or pursue additional education). They can involve an explicit choice of task or strategy (hunting and pecking or touch typing), a choice of more general factors, such as how much mental effort to allocate to a specific task (skimming or perusing a textbook), or implicit processes that are implemented without explicit planning, intent, or even awareness. In this study, we focus on the latter. Although the specific mechanisms responsible for assessing the relative costs and benefits of investing in learning may differ, the factors (costs vs. benefits of investing time and effort) and principles (optimization of the tradeoff between the near- and long-term benefits) that govern the choice are likely to be the same. People do learn to touch type, follow new recipes, and pursue additional education, which suggests that a “value of learning” shapes how we allocate cognitive control.

While current theories admit that cognitive control allocation takes into account the future discounted value of control-demanding activities, explaining why people may invest in tasks that may take time and effort (e.g., refs. 57), these have largely assumed that task proficiency is a known and stable quantity. In doing so, they have not generally taken into account the changes in proficiency that can occur with learning, and therefore the corresponding changes in the expected value of allocating control to such tasks (8). The failure to account for such effects poses a challenge, that has been described as the cognitive “effort paradox” (9), which questions why people should rationally choose to invest in performing tasks (such as attempting to touch type) that are more effortful and less rewarding than immediately available alternatives (such as the hunt and peck strategy). Here, following recent empirical work in rodents (10), we address this dilemma by incorporating “learnability” into the calculation of future value and corresponding allocation of control, and ask whether humans consider the value of learning in allocating cognitive control.

There has been, to our knowledge, no direct test of whether the value of learning plays a part in how humans allocate cognitive control, but several lines of evidence suggest that the reward implications of learning do indeed influence the allocation of cognitive control in primates, including humans. For instance, human infants and macaques preferentially attend to stimuli that are of moderate complexity and intermediately surprising (1113). Meanwhile, adults choose their curricula according to similar principles—those of self-directed learning (14), and of selecting tasks of intermediate difficulty (15, 16)—in order to maximize their learning and reward. These results are in line with normative theories of learning that suggest that agents should favor items of intermediate complexity to most effectively fill gaps in their knowledge (1721): while items of low complexity may require less control, the fact that they are easy limits the learning they can confer (think of carrying out single digit addition). Likewise, while items of high complexity will require more control, their complexity may actually inhibit efficient learning because, for example, feedback may not be readily interpretable (think of asking questions in a language with limited proficiency).

Similarly, accounts of flow (22, 23) suggest that matching one’s skill and difficulty in a task can be subjectively rewarding because it can lead to maximal learning (24, 25). The connection here lies in the fact that although some learning can occur when the task is easy and the agent is skillful (or the task is hard and the agent is unskillful), the most efficient learning will occur when the task is easy if the agent is a novice, and hard if the agent is an expert. This way, the agent’s control investment in learning will be most effective. As with flow, recent accounts of boredom (15) and fatigue (26, 27) suggest that these subjective experiences promote the reallocation of cognitive resources in the service of more effective learning. The present study directly empirically tests the extent to which the value of learning impacts the allocation of cognitive control in humans. We begin by formalizing a model of how learning can be incorporated into the computation of the expected value of control (6), and then present the results of an empirical study designed to test this model.

Model

EVCL-LDDM.

To generate predictions about how people may allocate control as a result of expected learning, we integrated two previous models of decision making, learning, and control. The learning drift-diffusion model [LDDM; (10)] addresses learning and decision making, while the EVCL model [EVCL; (28)] addresses how optimization of control can consider the effects of learning. The resulting EVCL-LDDM model (Fig. 1) is composed of three components: 1) a processing mechanism to make choices; 2) a learning mechanism to improve choices through practice; and 3) a control mechanism responsible for optimizing performance according to some criterion, such as maximizing cumulative reward. Critically, this optimization takes account of the anticipated effects of learning. In the following two sections, we describe how these components are implemented by combining the LDDM and EVCL models.

Fig. 1.

Fig. 1.

EVCL-LDDM model, study rationale, and design. Model contains: (A) a cognitive control mechanism that evaluates the expected value of control for learning (EVCL; see text for explanation) and allocates control over the decision making process; (B) an integrative decision making mechanism that implements a standard drift diffusion model (DDM) in the form of a simple recurrent linear neural network; and (C) an error-corrective learning mechanism (LDDM). The control mechanism allocates control by modifying the evidence accumulation threshold of the DDM across trials, with higher thresholds associated with greater control (more accumulation). If the threshold is set to maximize cumulative reward over some time horizon, it is equivalent to an optimal EVCL model that provides a normative account for the effects of learning, using threshold as the cognitive control variable. Learning occurs by adjusting the network weight across trials, which is equivalent to adjusting the attentional component of the drift rate of the DDM. (D) The EVCL-LDDM model predicts that skill level and learning rate (learning expectations) are used to determine the optimal threshold (control allocation) (Fig. 4). (E) Thus, we hypothesize that a higher starting drift rate and change in drift rate (corresponding fitted variables) during an inaugural learning experience (Block 1, see F) should predict a higher threshold (control allocation) and decision time at the start of a subsequent learning experience (Block 2, see F). (F) To test our hypothesis, we designed a dot-motion experiment with a first block to induce learning expectations, and a second block to measure the effect of those expectations on control. In the model, skill level and learning rate are used to simulate forward in time to compute the optimal control allocation—optimal threshold of the DDM—under some policy, such as maximizing cumulative reward. Because the model is forward-looking, skill level and learning rate are effectively priors on learning expectations. These learning expectations were induced in participants by undergoing a learning experience (Block 1) and measured by fitting their starting drift rates and changes in drift rate. The effects of these learning expectations on control were measured by regressing threshold and decision time at the start of a subsequent learning experience (Block 2). Because the stimulus only contained noise in Block 2, the thresholds and decision times set by participants reflected their learning expectations from Block 1. (G) Summary of the variable correspondences between the theory, and the empirical variables.

LDDM.

The LDDM builds on the standard DDM (29) of decision making, in which an agent accumulates evidence favoring each of two options, and makes a choice when the evidence for one of these crosses a specified threshold. This procedure can be used to model both response times (number of steps of integration at which a threshold is crossed) and accuracy (whether the threshold for the correct response was crossed).

In its simplest form, the DDM is fit to empirical data by assuming that its parameters—and, in particular, the drift rate (corresponding to signal efficacy) and decision threshold—are fixed across trials (3033). Considerable work has addressed more dynamic versions of the model in which parameters vary within a trial, such as changes in drift rate as a function of attention (34, 35) or urgency (36) and/or a progressive reduction in threshold to impose an upper limit on response time (3739). Some work has addressed changes that may occur trial-to-trial (such as sequential adjustment effects in response to errors or conflict) (40).

Nevertheless, considerably less work has addressed changes in parameters that may occur over the course of learning. For example, for a stimulus of constant difficulty, an increase in the drift rate across trials can reflect an improvement in the agent’s skill level. Thus, learning can be represented as an increase in drift rate. The LDDM takes this approach, implementing a DDM in the form of a simple linear recurrent neural network model, in which changes in drift rate occur through changes in connection strengths as a function of a SE-driven (backpropagation) learning algorithm (Fig. 1 B and C). Furthermore, the model has an analytical solution for its average learning dynamics [Eq. 6; (10)]. With this solution, the threshold (and resulting average learning trajectory) that optimizes a specified objective, such as maximizing cumulative reward over a reasonable time horizon, can be easily determined without computationally intensive simulations. Learning trajectories and their reward outcomes simulated under different objectives can then be compared to empirically measured learning trajectories in order to determine which objective participants may most likely have employed.

EVCL.

In the DDM (and LDDM), threshold determines how much evidence integration takes place. This in turn governs the speed–accuracy tradeoff: low thresholds favor speed over accuracy; high thresholds favor accuracy over speed (Fig. 2 A and B). Previous work has shown that, for a fixed set of other parameters, there is a single optimal threshold that balances speed and accuracy to maximize current reward rate (41) (e.g., low threshold in Fig. 2E). If one of these parameters changes (e.g., drift rate changes through learning), then the threshold must be adapted to continue to optimize performance. However, what has been less well recognized—but may be equally important—is that threshold can also directly impact learning.

Fig. 2.

Fig. 2.

Explanation of optimal control (threshold) computation. In the EVCL-LDDM model, a greater threshold can enhance learning, as reflected in a greater increase in signal-to-noise ratio (SNR) over training. However, the current reward rate and cumulative reward do not necessarily increase monotonically with threshold. In order to compute the optimal threshold for some policy (i.e., maximize total cumulative reward) for a given set of parameters (skill level and learning rate), we simulate the model with different thresholds, and visualize the (A) decision time, (B) error rate, (C) SNR (drift/noise), and (D) change in SNR (dSNR/dt). We then compute the optimal threshold according to two policies, (E) a greedy policy that seeks to maximize reward rate, and (F) an EVCL policy that seeks to maximize cumulative reward.

Fig. 2 illustrates the effect that choice of threshold has on learning and reward trajectories. For a particular skill level and learning rate, a threshold that is higher than the current optimal value (low threshold)—that is, one that overly favors accuracy over speed (intermediate or high thresholds in Fig. 2 A and B)—may nevertheless provide more accurate information that can be exploited by learning, serving to increase drift rate more rapidly over subsequent trials (Fig. 2 C and D, where SNR is drift/noise). A swifter increase in drift rate may lead to higher reward rates in the longer term that compensate for lower immediate reward rates (Fig. 2E). However, the most learning does not necessarily mean the most reward: Reward is accrued over time, and higher thresholds require more time between rewards. Integrating over reward rates reveals that—for this particular skill level and learning rate—an intermediate threshold best balances learning and reward to maximize cumulative reward (Fig. 2F). This normative process can be formalized by combining the LDDM with the EVCL model.

In EVCL, the agent uses estimates of its current level of skill (in this setting, the agent’s internal component of drift rate) and learning rate (approximated empirically as the change in drift rate over time) to determine how to allocate control (adjust threshold), by estimating how this will impact its objective (e.g., maximize future cumulative reward). Specifically, it weighs the costs of a particular control signal setting (i.e., a higher threshold will lead to longer response times and thus lower reward rates in the near term) against the longer term benefits (increases in drift rate that will yield increases in reward rate over time), and adjusts the control signal (threshold) based on these estimates in order to maximize the amount of reward predicted over a given time horizon (e.g., number of trials).

The relative value placed on immediate and future reward can be modulated with a discount factor on future reward. An agent that optimizes control based only on estimates of current reward rate, effectively ignoring the effects of control on learning and future expected reward, follows a “greedy” policy, the objective assumed in traditional models of control allocation. An agent that values both immediate and future reward, and accounts for the value of learning, follows an EVCL policy. In this study, we seek to test whether humans produce behavior that better adheres to a greedy objective that does not consider future learning in its control allocation, or to an EVCL objective that does.

Results

Task.

We hypothesized that participant experience with a perceptual task would generate learning expectations (i.e., skill level and learning rate, estimated using starting drift rate and change in drift rate, respectively) that should influence the allocation of control (threshold setting) at the outset of a sufficiently similar new task (Fig. 1E).

To test this hypothesis, we employed a simple psychophysical task extensively used to study decision making and learning (4244) that requires participants to identify the predominant direction in which a field of randomly placed dots are moving, and in which task difficulty can be manipulated by varying the number of dots moving coherently vs. randomly. Participants were tested in two blocks (Fig. 1F). During Block 1, participants encountered difficult but learnable stimuli, allowing them to generate expectations about the learnability of a similar task. To create variance in participants’ experience of the task during Block 1 that we could later exploit in our analyses, participants were randomly assigned to 5, 10, or 15% coherence conditions, ranging in demand from harder to easier. During Block 2, participants’ encountered a similar task with slightly different stimuli. Unbeknownst to participants, the new stimuli were no longer learnable, allowing us to measure, in as unbiased a way as possible, the influence that participants’ experience (starting drift rate and change in drift rate) in Block 1 had on their control allocation (threshold setting) at the outset of Block 2.

Learning Scales with Stimulus Difficulty.

To evaluate the effectiveness of the manipulations meant to vary learning expectations, we computed mean error rates and decision times in Block 1 for every coherence condition (Fig. 3A). Linear regressions (Eqs. 16 and 17) confirmed that average error rate and decision time decreased over trials, indicative of learning, and that the degree of decrease corresponded to stimulus difficulty (SI Appendix, Tables S1 and S2). Group-level posterior estimates from hierarchical linear Bayesian DDM fits confirmed our raw behavioral data analysis and revealed that, as expected, both starting drift rate and change in drift rate on average increased with decreasing stimulus difficulty in Block 1 (Fig. 3E). As such, they indicate that our experimental objective of inducing a learning prior based on learning outcomes was effective.

Fig. 3.

Fig. 3.

Performance in Block 1 predicts control allocation in Block 2. (A) Decrease in mean error rates and decision times (25 trial bins, 95% CI) across motion coherence conditions in Block 1 (5%, n = 58; 10%, n = 50; 15%, n = 51) indicate learning scales with stimulus difficulty. (B) Higher motion coherence in Block 1 leads to higher mean decision time during first 25 trials of Block 2. Decision time can be used as a proxy for threshold and therefore control because drift rate is 0 (coherence is 0%) during Block 2 (SI Appendix, Fig. S9A, Top Left panel). (C) Model error rate and decision time for qualitative fits to 5, 10, and 15% motion coherence conditions in Block 1. (D) Model optimal decision times using qualitative parameter fits to Block 1 motion coherence data from EVCL policy (Right panel), and not greedy policy (Left panel), match experimental results in B. (E) Hierarchical linear Bayesian DDM regression fits to Block 1 show that starting drift rate (skill level) and change in drift rate (learning rate) increased as functions of coherence condition, confirming that participant learning on average scaled with stimulus difficulty (as indicated in A). (F) Starting initial threshold was similar across coherence conditions, indicating learning priors did not differ at the outset of the experiment. Threshold generally decreased over trials.

To verify that participants’ learning expectations at the outset of the experiment did not differ substantially across conditions, we examined group-level posterior estimates for initial threshold in Block 1, and found that these largely overlapped across coherence (Fig. 3F, Left panel). Thresholds decreased over Block 1, with a trend toward a larger decrease for the easiest 15% coherence condition (Fig. 3F, Right panel). It is not uncommon for thresholds to decrease over the course of a simple perceptual task (e.g., ref. 45), potentially reflecting factors such as boredom (46) and fatigue (47), which we do not consider here.

Previous Performance Predicts Subsequent Control Allocation.

Next, we tested whether learning expectations based on experience in Block 1 shaped control allocation at the outset of Block 2. Examination of the raw data revealed that at the start of Block 2, mean decision time scaled with coherence condition during Block 1 (Fig. 3B). This observation suggested that participants chose larger thresholds for conditions in which they experienced larger improvements (i.e., greater learning) in the previous block.

As a first test of our model, we simulated optimal decision times for Block 2 based on participant performance in Block 1. We first qualitatively fit the mean error rate and decision time data from Block 1 (Fig. 3C). Then, using these parameters, we simulated optimal decision times under greedy and EVCL policies. We found that predictions of the EVCL policy, maximizing cumulative reward, better fit the participants’ mean decision times at the start of Block 2 (Fig. 3 B and D).

Model Predicts Learning Expectations Determine Optimal Control Allocation.

Our analysis of mean performance suggested that participants considered learning expectations when choosing their control allocation. In order to confirm and directly test this suggestion, we generated optimal thresholds from the greedy and EVCL policies across a range of skill levels and learning rates (corresponding to the range of starting drift rates and changes in drift rate fit to participants during Block 1). The greedy policy prescribes that the optimal threshold should increase with starting drift rate (skill level) but, by construction, not change in drift rate (learning rate) (Fig. 4A). In contrast, the EVCL policy prescribes that optimal threshold should increase as a function of both starting drift rate (skill level) and change in drift rate (learning rate) (Fig. 4B). The EVCL policy additionally reveals an effect associated with the interaction of skill level and learning rate, such that the greater the skill level, the lesser the effect of learning rate because the agent should have less left to learn, an approximately negative linear interaction for the range of parameters explored (Fig. 4).

Fig. 4.

Fig. 4.

Learning expectations shape control allocation. To determine how learning expectations shape control allocation in the EVCL-LDDM, we computed optimal thresholds (Fig. 2) across learning rates (λ) and skill levels (u) for a greedy policy (maximizing current reward rate) and an EVCL policy (maximizing total cumulative reward), and plotted these as functions of the empirical variables starting drift rate, Au(trial=1), and change in drift rate, (Au(trial=200)Au(trial=1))/2. (A) With a greedy policy, optimal threshold depends only on starting drift rate. (B) With an EVCL policy, optimal threshold depends on both starting drift rate and change in drift rate, with a negative interaction (in the range of parameters qualitatively fit to Block 1) such that the higher the starting drift rate, the lesser the effect of the change in drift rate. (C) Regression predictions from the hierarchical linear Bayesian DDM fits to Block 2 (SI Appendix, Fig. S9) indicate that threshold depended on change in drift rate, starting drift rate, and that these effects had an approximately negative linear interaction, better matching the predictions of the EVCL policy in B. Dots in A and B are the model predictions, and lines are linear approximations. The color gradient indicates starting drift rate (Au).

Learning Expectations Shape Control Allocation.

We sought to quantitatively test our model’s predictions about how learning expectations (starting drift rate and change in drift rate) shape control allocation (threshold setting). To do so, we regressed threshold at the start of Block 2 on starting drift rate, change in drift rate and their interaction during Block 1 with a hierarchical linear Bayesian DDM regression (Eq. 14). We included an additional regressor to control for a possible confound in our design: elevated thresholds at the beginning of Block 2, rather than reflecting strategic allocation of control, could have been due simply to temporal autocorrelation of threshold setting between the end of Block 1 and the start of Block 2. This would have produced the observed results (Fig. 3B), even for the greedy policy, in that reward rate optimization should produce higher thresholds in Block 1 (in which all participants experienced coherences greater than 0% and thus some signal) than in Block 2 (in which all participants experience 0% coherence). Thus, thresholds in Block 2 higher than predicted by the greedy policy could reflect simply slow—rather than strategic—adaptation of threshold. To control for this possible temporal autocorrelation of threshold, we included the inferred final threshold from Block 1 as an additional regressor.

The regression revealed that Block 1 starting drift rate and change in drift rate both had positive effects on threshold at the outset of Block 2 (Fig. 4C, see SI Appendix, Fig. S9 for posteriors on coefficients). There was also a negative interaction between the two main effects, such that the effect of change in drift rate on threshold was smaller the larger the magnitude of the starting drift rate. These results, showing a sensitivity to both skill level and learning rate, are consistent with the predictions made by the EVCL threshold policy that maximizes cumulative reward (Fig. 4 A and B).

An additional analysis on decision time investigated whether participants’ learning expectations led them to engage in behavior that was costlier in terms of immediate reward rate. In the model, the cost of a larger control allocation (a higher threshold) comes in the form of a reduced reward rate. Although threshold is the cognitive control variable, it is decision time, the resulting behavior, that leads to a reduction in reward rate. A linear mixed effects regression of decision time at the start of Block 2 (Eq. 15 and SI Appendix, Fig. S10 and Table S6) revealed positive effects for starting drift rate (β=0.320, 95% CI[0.167,0.473], P<0.001) and change in drift rate (β=0.534, 95% CI[0.228,0.839], P=0.001) in Block 1, and a negative trend for the interaction between starting drift rate and change in drift rate on decision time at the start of Block 2 (β=0.539, 95% CI[1.190,0.112], P=0.107). These effects were above and beyond the positive effect of inferred final threshold in Block 1 (β=0.515, 95% CI[0.417,0.613], P<0.001). As with the regression on threshold, the decision time regression results are consistent with the predictions made by the EVCL-LDDM model equipped with an EVCL threshold policy that maximizes cumulative reward and takes account of possible future learning, and not with a greedy threshold policy that maximizes current reward and does not consider learning.

Discussion

The work presented in this article suggests that people selectively allocate cognitive control in response to learning expectations based on prior experience. We first introduced a model, the EVCL-LDDM, that combines mechanisms that model binary choice and learning [LDDM; (10)] with a model of strategic control allocation that takes learning into account [EVCL; (28), recently expanded and generalized in ref. 48). The EVCL-LDDM prescribes that an agent that aims to maximize total cumulative reward over some horizon should allocate cognitive resources in response to its learning expectations. In the present application, it predicts that an agent should modulate its deliberation time during evidence accumulation (by manipulating decision threshold) as a function of its priors on estimates of its skill level and learning rate, as well as their interaction. In contrast, an agent that seeks only to optimize current reward rate should choose its deliberation time based solely on its current skill level, as has been described previously (e.g., ref. 41). Decision threshold functions as a cognitive control variable in our setting because modulating threshold based on learning expectations requires strategically overriding the likely prepotent response of faster responding that leads to both less integration, and also higher current reward rates.

We found support for these predictions from an empirical study in which participants were given the opportunity to develop learning expectations through experience with a task in a first block of trials, and then evaluated on the impact of those learning expectations on their decision parameters in a second block of trials. As predicted by the EVCL-LDDM, participants’ skill level and learning rate predicted decision thresholds and deliberation times at the start of the second block. These effects could not be fully explained either by a greedy model, reflecting a common normative objective of current reward rate maximization, or by simple temporal autocorrelation of threshold setting.

These results are consistent with the hypothesis that humans allocate cognitive control in response to their learning expectations, and with the first formal model of how this process may occur. Our results suggest that people strategically allocate their mental resources (measured here as their willingness to deliberate) in order to modify their own cognitive bounds (measured here as their skill level). Although perhaps unsurprising, this process challenges an overlooked assumption in some of our best models of cognition: that (all) cognitive bounds are fixed (8). Whereas some cognitive bounds may be beyond our control to adjust (such as a bottleneck at the level of speech production), many, such as those related to skill, are malleable, and that malleability should be factored into the strategic allocation of cognitive resources.

Related Work.

The adaptive allocation of control based on learning expectations is related to other expressions of the human capacity for adaptive control. For example, one active line of work concerning the role of exploration in the context of reinforcement learning, has suggested that people exhibit greater exploratory behavior when there is a greater potential to improve future rewards, such as when time horizons are longer (49). In fact, people trade immediate reward (exploitation) for information (exploration) if it has the potential to improve future performance (21), an observation closely related to the impact of learning expectations on control allocation in our study. Similarly, work grounded in control theory has found that people strategically weigh costs and benefits when deciding whether to explore unknown actions (e.g., playing through an entirely new song is better suited for a rehearsal than a recital) (50). Such findings are in line with the theory we have tested here: people’s decisions to allocate cognitive control while accounting for learning are sensitive to the balance between its costs (in terms of time and effort) and benefits (in terms of opportunity and impact on future reward).

These effects also closely relate to the commonly observed nonmonotonic relationship between task difficulty and engagement: people are generally less engaged by tasks that are either too easy (e.g., boredom) or difficult (e.g., frustration), and most engaged at intermediate levels of difficulty. Work within the context of reinforcement learning has suggested that optimal levels of engagement (usually observed when performance is in the range of 70 to 85%) closely correspond to the highest learning rate (24). The idea of an optimal level of engagement undergirds work that explicitly addresses the phenomena of boredom, fatigue, and the state of “flow.” Boredom may reflect an internal signal indicating that the information value of the current task is low and that exploration (i.e., task disengagement and information seeking) may hold greater value for future reward (15). Fatigue may reflect the relatively higher value for learning of deferring immediate action in favor of the internal replay of recent events (i.e., favoring model-based vs. model-free learning; refs. 26 and 27). Conversely, it has been suggested that the mental state of “flow” may reflect an optimum in the level of information acquisition and learning in the current task (22, 24, 26). Flow has recently been formulated as the mutual information between an agent’s desired states and its ability to reach those states (23) and implemented in a computational model of decision making (in which choices are based on information-seeking) that shows that maximizing flow leads to faster learning than a policy based on reward-maximization (25). Moreover, changing how a task (shooting a basketball) is framed (make at least 1 out of every 10 shots, or at least 10 shots in a row) can dramatically change the experience of flow, and consequently people’s engagement in the task based on their current level of skill (51). Changing the task framing is tantamount to changing how reward is represented, and consequently what the expected value of control is, providing not only a potential explanation for people’s willingness to engage in a task (based on their current level of skill and the task’s current framing) but also a normative reason for the experience of flow.

All of these lines of inquiry share with the EVCL-LDDM model presented here the idea that people allocate control by seeking to optimally balance immediate and long term opportunities for reward and that information gathering and learning are critical factors in achieving the latter. A better understanding of how these various formulations relate to one another, and an integrated account of how people balance this tradeoff is an important direction for future research.

Limitations and Future Work.

One important limitation of the present work, and the related lines of work outlined above, is that they all focus on relatively simple tasks executed over relatively short time horizons. While this has proven valuable for the construction of formally rigorous theory and the ability to test it empirically, future work will need to consider the extent to which the forms of estimation and optimization assumed in such models scale to more complex tasks, and over longer time horizons.

The algorithms and/or heuristics people use to optimize performance present an important and particularly interesting question (52, 53). We implemented the EVCL-LDDM model using dense optimization, a process in which all combinations of parameters within a reasonable range are tested in order to find the best combination. This process does not scale tractably, and thus is not likely to reflect how people actually approach the problem. Another approach would be to assume that agents use Bayesian inference to reduce uncertainty over time (5457). This could be tested by manipulating the uncertainty of learning trajectories (e.g., by occasionally withholding feedback) to explicitly examine the process through which people may form their learning expectation priors (58). Beyond previous task performance, participants may use other factors, such as the similarity of a new situation to a previous situation, to estimate the value of control in the context of learning. This similarity comparison could also involve Bayesian estimation, here of features that are most predictive of the value of control (e.g., refs. 59 and 60), and/or the use of episodic memory to retrieve prior similar circumstances and the associated utility of control (e.g., ref. 61). One can reasonably imagine a mechanism for constructing a predicted learning curve for a new task by probing episodic memory for previously encountered tasks and learning outcomes. Such a mechanism would circumvent the implausible nature of the dense optimization in the normative approach applied here. Moreover, in our model, control functioned by determining participants’ willingness to deliberate, but previous work has also modeled control through increased attention to the stimulus (cf., ref. 62). Future elaborations of our model could incorporate such attentional mechanisms in order to better investigate how people strategically allocate control in the context of learning.

Finally, understanding individual differences in skill, learning rate, and how these impact optimization in the allocation of control is another valuable direction for future work. In the present study we exploited such intersubject variability to test the relationship between learning and control. However, this may have limited our sensitivity to fully evaluate this relationship (e.g., due to overlap in the starting drift rates and changes in drift rate over our coherence conditions; see SI Appendix, Fig. S8). Future work could directly manipulate the rate of learning (e.g., by selecting a task with less intersubject variability and introducing more difficulty conditions). Furthermore, a deeper understanding of individual differences in skill, learning, and control allocation themselves may be important for understanding how people differ under different task settings and demands, and how this may interact with patterns of performance associated with various clinical conditions associated with disturbances of attention and/or motivation.

Materials and Methods

EVCL-LDDM.

LDDM.

For a complete description of the LDDM, please see ref. 10. In brief, the decision variable y^ receives the sum of the scalar stimulus input x multiplied by a scalar learnable weight u and ηN(0,co2), the output (or accumulation) noise.*

y^(t+dt)=y^(t)+u(trial)x(t)+η(t). [1]

The stimulus input is the scalar Gaussian x(t)N(Aydt,ci2dt), where A is the signal strength, y is the true stimulus identity y(trial)=±1, and ci2 is the irreducible stimulus noise. The stimulus input is scaled by a modifiable input weight u(trial) that is adjusted using an error-driven learning algorithm. Once y^ reaches a predetermined threshold, z, a choice is made (Fig. 1B). Variables that update every trial are a function of “trial,” whereas variables that update within trial during the evidence accumulation are a function of t.

The state of the output layer y^ and the evidence threshold z at which a choice is triggered are equivalent to the accumulation variable and threshold in a standard DDM, respectively (Fig. 1B, dotted gray lines). Accordingly, the network variables can be reexpressed in the same way as the parameters of a standard DDM, in terms of the drift Au, the drift-to-noise ratio A¯ (i.e., SNR) and threshold-to-drift ratio z¯.

A¯=A2u2u2ci2+co2, [2]
z¯=zAu, [3]

which, following ref. 41, allows us to estimate the mean error rate (ER) and decision time (where decision time is the difference between reaction time and nondecision time DT=RTt0) for these parameters.

ER=11+e2z¯A¯, [4]
DT=z¯tanh(z¯A¯). [5]

LDDM has the advantage of being able to adjust its perceptual weight through error-driven learning (Fig. 1C), allowing the model to conduct gradient descent on the hinge loss, Loss(trial)=max(0,1y(trial)y^(trial)), (63). The hinge loss, popular in binary classification tasks, updates the scalar weight after both correct and incorrect responses with a small learning rate λ, according to u(trial+1)=u(trial)λδLoss(trial)δu. The process of adjusting the strength of the perceptual weight in LDDM in response to feedback is tantamount to modifying the internal component of the drift rate in a traditional DDM. Insofar as the drift rate for a stimulus of fixed strength is commonly taken to index an agent’s familiarity with and/or skill in processing the stimulus, the effects of error-driven learning can be used to model learning of the task through experience.

LDDM has the benefit of a reduction that results in an analytical solution to the average learning dynamics of the network. To arrive at this reduction, we assume that the learning rate is small (λ1) and that the weight changes little on any given trial, such that the gradient dynamics are driven by the mean update u(trial+dt)=u(trial)λδLoss(trial)δu. This solution allows us to easily determine reward outcomes based on different objectives or policies governing the choice of thresholds and to compare those trajectories to data (10).

τ~ddtA¯(t)=2A¯(t)A¯c1A¯(t)A¯5/2ER(t)DT(t)+Dtot(t)DT(t)log(1/ER(t)1)A¯1A¯(t)A¯2. [6]

The learning dynamics depend on the learning rate in the network (expressed through the time constant τ˜, related to the learning rate in the network), the maximum skill level possible for the task (the asymptotic achievable SNR A¯=A2/ci2 in the limit of very large perceptual weights), the agent’s skill level at the beginning of the task (the inaugural SNR A¯(0)), the ratio of output to input noise (cco2/ci2), and the dynamic choice of threshold (z(t)) throughout the task. The dynamic choice of threshold comes through implicitly through the changing ER(t), DT(t) and Dtot(t), where Dtot(t)=(1ER(t))Dcorr+ER(t)Derr is the average nondecision task engagement time per trial. Dcorr and Derr are the average nondecision task engagement times after correct or incorrect choices. Each is the sum of the response-to-stimulus (RSI) interval after a correct or incorrect response, D~corr or D~err, and the nondecision time component of reaction time t0, e.g. Dcorr=D~corr+t0.

EVCL model.

Following ref. 28, in EVCL, the expected value of control for a candidate control signal in a particular state at a particular timestep, trial in this study, is given by the difference between the expected payoff of control in that state and the cost of that candidate control signal.

EVCtrial(signal,state)=E[Payoff(signal,state)]Cost(signal). [7]

The signal and state are theoretical variables that must be grounded to a specific task at hand. Here, the agent’s candidate control signal is composed of its threshold (z), whereas the agent’s state is composed of its current skill level (expressed as either drift rate Au or SNR A¯), asymptotic achievable skill level (A¯), learning rate (λ), and the number of trials completed thus far.

The expected payoff is computed as the expected value of the outcome given the candidate signal and state. The outcome is the expected new state (given the candidate control signal and current state), for which all elements remain the same except for the number of completed trials and the current skill level, which is updated via the average LDDM dynamics (Eq. 6). The expression of the payoff here as an expectation reflects the fact that in hypothetical simulations of the LDDM network, the expected outcome would be the average outcome across simulations. Rather than simulate the network, we directly compute its average dynamics (Eq. 6), and therefore expected outcome and resulting payoff.

E[Payoff(signal,state)]=iP(outcomei|signal,state)·Value(outcomei). [8]

The value function is, in turn, composed of two elements.

Value(outcome)=R0(outcome)+γ·maxj[EVCtrial+1(signalj,outcome)]. [9]

The first, R0, is the immediate reward for the outcome at the current trial. The second, γ·maxj[EVCtrial+1] (where j indexes the possible control signals, i.e. possible evidence thresholds, at the next timestep trial+1), is the future discounted reward for the candidate control signal that yields the greatest expected reward when the outcome of the current state is used as the next state. The discount factor γ controls whether the model is fully myopic (γ=0) or forward-looking with no discounting (γ=1) (SI Appendix, Fig. S19).

Immediate reward.

We define the immediate reward R0 (Eq. 9) as the instantaneous or current reward rate, iRR. R0 is a function of the current outcome Au, which, with the candidate control signal z, can be used to calculate mean ER and DT (Eqs. 4 and 5), which can in turn be used to compute expected immediate reward rate.

R0iRR=1ERqERDT+Dtot, [10]

where Dtot=RSI+t0 captures the response-to-stimulus interval and the nondecision component of reaction time t0. Finally, we allow for q, a reward penalty for errors referred to as an accuracy bias (6466). We included this extra parameter because we found that empirical decision times were generally higher than we could qualitatively fit with reasonable parameters when considering a pure reward rate (q=0) formulation. Further, because the cost of a high threshold is implicitly included in the agent’s reward rate (a higher threshold would lead to a higher average DT, and thus a lower reward rate), we decided to omit the explicit control signal cost term in Eq. 7. This explicit cost term could reflect the cognitive cost of changing thresholds, but this is probably negligible compared to the cost in reward rate of a particular threshold setting.

Optimal control signal.

Because the value function of the current state uses the expected outcome of the current state to compute the EVC for the next state (Eq. 9), EVCL can be used to recursively simulate the consequences of its control choices over a reasonable range of control signals (threshold choices) and up to some tractable future time horizon. The optimal control signal signal (threshold choice) for the current state is then chosen by selecting the control signal with the maximum EVC.

signalargmaxi[EVC(signali,state)]. [11]

Threshold control policies.

As a benchmark, we compare the EVCL policy that maximizes the integral of reward over some predetermined time window with a simpler normative greedy policy that maximizes instantaneous or current reward rate (67). The greedy policy produces behavior that is equivalent to the theoretical solution for the speed–accuracy tradeoff, termed the optimal performance curve (SI Appendix, Fig. S16) (41).

To toggle between the greedy and EVCL policies, we manipulate the discount factor. For the greedy policy, we make the model fully myopic (set the discount factor γ=0 in Eq. 9) in order to have it choose the threshold that maximizes current reward rate. For the EVCL policy, we make the model value rewards at the end of the horizon as much as immediate rewards (set the discount factor γ=1), and thus have it choose the threshold that maximizes total cumulative reward. A discount factor in between effectively shortens the horizon of the EVCL policy (SI Appendix, Fig. S19).

Although the greedy policy is fully myopic, it is still responsive to previous experience (such as that of Block 1) in the form of its predicted current skill level (SI Appendix, Fig. S17). Although the EVCL policy is forward-looking, it will only differ from the greedy policy when there are predicted future dynamics (such as through expected learning) that it can exploit (SI Appendix, Fig. S20). We further include in the supplement a “sophisticated” greedy policy that is free to choose any threshold in the first trial (before subsequent thresholds must maximize current reward rate) that can help it maximize total cumulative reward (SI Appendix, Fig. S18). This model did not produce predictions that matched the empirical results.

For tractability in threshold optimization, while forward looking simulations sampled different possible thresholds, threshold was kept fixed over the duration of each simulation (i.e., over its temporal horizon). Finally, because we are only interested in evaluating the impact of learning expectations on initial threshold setting, we report the optimal threshold at the beginning of the horizon.

Study.

Design.

The study was composed of a baseline signal detection task (SI Appendix) during which we measured participants’ nondecision times (t0) in order to aid with DDM fits (SI Appendix, Fig. S3) followed by two experimental task blocks of equal length (200 trials) of random dot motion (Fig. 1F). During Block 1, the inducement block, participants were presented with motion coherences of 5, 10, or 15%. Trials started with a fixation cross, followed by the stimulus and feedback (SI Appendix, Fig. S1). A pilot study that included confidence judgments and surveys (SI Appendix) indicated that in this coherence range participants reliably learned but were not reliably aware of their learning (SI Appendix, Fig. S2). Operating just beyond participants’ awareness of learning would allow us, we reasoned, to measure control allocation while reducing the interference of overt strategies. During Block 2, the measurement block, participants were presented—unawares—with a motion coherence of 0%, i.e., random noise, in a direction orthogonal to what they saw in Block 1. The change in motion direction and the equal block length (despite the fact that we would only consider early trials in our analysis) were chosen to communicate that Block 2 contained a very similar but distinct task to Block 1. Presenting participants with random noise in Block 2 served a twofold purpose. First, it would allow us to measure participants’ choice of early thresholds (i.e., first 25 trials) based on—we hypothesized—learning expectations (or priors) formed during Block 1 before they were substantially updated with new evidence from Block 2. Second, it would allow us to use decision time as a secondary measure of control allocation: a motion coherence of 0% all but guarantees a drift rate of 0, which means that decision time depends entirely on threshold choice. We chose 25 trials as the window for analysis during Block 2 in an attempt to balance our experimental desire for few early trials with the reliable and accurate recovery of latent participant parameters. Post hoc analyses revealed that the results were robust to the window of trials analyzed (SI Appendix, Fig. S12)

Fig. 1G includes a summary of our theoretical and empirically measured and fitted variable correspondences. SI Appendix, Table S13 describes terms used in the article.

Participants.

We collected data online from 197 participants, each receiving US $4.80 USD ($10.77 per hour), using Prolific (https://prolific.co). Participants provided written informed consent. The study was approved by the Princeton University Institutional Review Board. After basic engagement exclusions (SI Appendix, Fig. S5), 159 participants remained. Of these, 58, 50, and 51 performed the 5, 10, and 15% coherence conditions during Block 1.

DDM regressions.

We performed hierarchical linear Bayesian DDM regression fits on the data. During Block 1, we regressed drift rate and threshold on trial to estimate their evolution over the block.

drift rate1+trial+(1+trial|participant), [12]
threshold1+trial+(1+trial|participant). [13]

Participant drift rate intercepts (starting drift rate) and slopes (change in drift rate) were used as approximate measures of skill level and learning rate.

During the first 25 trials of Block 2, we regressed threshold on starting drift rate, change in drift rate, their interaction, and inferred final threshold from Block 1.

threshold1+starting drift ratechange in drift rate+inferred final threshold+(1|participant). [14]

To account for and measure effects above and beyond the potential autocorrelation of threshold in Block 1 and 2, we computed an inferred final threshold for each participant. Inferred final threshold was the Block 1 regression (Eq. 13) estimate of threshold after 200 trials based on individual intercepts and slopes.

Previous work has found that when data generated from a collapsing threshold model are fit with a fixed threshold model, factors that influence drift rate can erroneously appear to also influence threshold (68). We thus fit Block 1 with a linearly collapsing bound DDM in order to ensure accurate estimation of the effect of trial on drift and threshold. In Block 2, however, we expected (and verified) drift rate to be 0 (SI Appendix, Fig. S9C), eliminating the concern of incorrectly estimating factors affecting threshold. We therefore considered it most parsimonious to base our principal conclusions, found in the main text, on fits from a standard DDM, the model upon which our theoretical model is based (SI Appendix, Fig. S8A). For completeness, we nevertheless found consistent results when fitting Block 2 with a linearly collapsing bound DDM (SI Appendix, Figs. S13–S15).

Linear mixed effects regressions.

logDT1+starting drift ratechange in drift rate+inferred final threshold+(1|participant). [15]

As a sanity check, we verified a coarse signature of learning by regressing decision time and outcome (correct, incorrect) on trial and coherence condition during Block 1 (SI Appendix, Tables S1 and S2).

outcome1+trialcoherence+(1|participant), [16]
DT1+trialcoherence+(1|participant). [17]

Qualitative model fits.

In order to find parameter ranges in our model that would match the data, we hand-tuned parameters to qualitatively match the mean error rate and decision time trajectories across coherence conditions during Block 1 (Fig. 3 A and C). We report these parameter values in Table S8. To compute the optimal decision times given the learning trajectories in Block 1 (Fig. 3D), we found the optimal thresholds for these qualitative parameters and computed decision times with Eq. 5. To generate valid predictions given our empirical data (Fig. 4), we computed the optimal thresholds according to a greedy or EVCL policy with a range of learning rates λ and skill levels u(0), as well as a higher SNR ceiling, set by the signal strength A, to span the values of starting drift rate and change in drift rate measured by the regression of threshold during Block 2 (Eq. 14 and SI Appendix, Table S9).

We note that participants’ decision times were too slow to be qualitatively matched with reasonable parameters when attempting to maximize either objective using simple reward rate (e.g. mean accuracy over mean time per trial, Eq. 10 where q=0). We found that we could match the range of participants’ decision times when we included an error penalty in the reward function [Eq. 10 where q>0; (41, 6466)]. Importantly, the parameter determining the error penalty magnitude q was the same for both the EVCL and greedy objectives.

Supplementary Material

Appendix 01 (PDF)

pnas.2416720122.sapp.pdf (56.1MB, pdf)

Acknowledgments

We would like to thank Harrison Ritz for discussions on drift-diffusion model fitting, and David E. Melnikoff for comments on earlier drafts. J.A.M.O. was supported by the Presidential Postdoctoral Research Fellowship at Princeton University, by the NIH Institutional Training Grant T32MH065214, and the Swartz Foundation Fellowship for Theory in Neuroscience at Princeton University. S.M. was supported by the Schmidt Science Fellows, in partnership with the Rhodes Trust, and the Carney BRAINSTORM program at Brown University. J.D.C. was supported by a Vannevar Bush Faculty Fellowship administered by the Office of Naval Research.

Author contributions

J.A.M.O., S.M., and J.D.C. designed research; J.A.M.O. performed research; J.A.M.O. contributed new reagents/analytic tools; J.A.M.O. analyzed data; J.A.M.O., S.M., and J.D.C. edited the paper; and J.A.M.O. wrote the paper.

Competing interests

The authors declare no competing interest.

Footnotes

This article is a PNAS Direct Submission.

*The accumulation or output noise η is injected directly into the accumulator y^ and reflects noise present in perceptual processing that is distinct from the irreducible input noise present in the stimulus.

The signal strength is equivalent to the absolute value of the difference between the possible stimulus means A=|μ1μ2|.

Data, Materials, and Software Availability

Data and code are available at https://osf.io/w46ev/ (69). A summary of the software used can be found in SI Appendix.

Supporting Information

References

  • 1.Logan G. D., Ulrich J. E., Lindsey D. R., Different (key) strokes for different folks: How standard and nonstandard typists balance Fitts’ law and Hick’s law. J. Exp. Psychol. Hum. Percept. Perform. 42, 2084 (2016). [DOI] [PubMed] [Google Scholar]
  • 2.Rumelhart D. E., Norman D. A., Simulating a skilled typist: A study of skilled cognitive-motor performance. Cogn. Sci. 6, 1–36 (1982). [Google Scholar]
  • 3.Logan G. D., Crump M. J., Hierarchical control of cognitive processes: The case for skilled typewriting. Psychol. Learn. Motiv. 54, 1–27 (2011). [Google Scholar]
  • 4.Logan G. D., Crump M. J., The left hand doesn’t know what the right hand is doing: The disruptive effects of attention to the hands in skilled typewriting. Psychol. Sci. 20, 1296–1300 (2009). [DOI] [PubMed] [Google Scholar]
  • 5.Kurzban R., Duckworth A., Kable J. W., Myers J., An opportunity cost model of subjective effort and task performance. Behav. Brain Sci. 36, 661–679 (2013). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Shenhav A., Botvinick M. M., Cohen J. D., The expected value of control: An integrative theory of anterior cingulate cortex function. Neuron 79, 217–240 (2013). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.Shenhav A., et al. , Toward a rational and mechanistic account of mental effort. Annu. Rev. Neurosci. 40, 99–124 (2017). [DOI] [PubMed] [Google Scholar]
  • 8.Musslick S., Masís J., Pushing the bounds of bounded optimality and rationality. Cogn. Sci. 47, e13259 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.Inzlicht M., Shenhav A., Olivola C. Y., The effort paradox: Effort is both costly and valued. Trends Cogn. Sci. 22, 337–349 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.Masís J., Chapman T., Rhee J. Y., Cox D. D., Saxe A. M., Strategically managing learning during perceptual decision making. eLife 12, e64978 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11.Kidd C., Piantadosi S. T., Aslin R. N., The goldilocks effect: Human infants allocate attention to visual sequences that are neither too simple nor too complex. PLoS ONE 7, e36399 (2012). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Cubit L. S., Canale R., Handsman R., Kidd C., Bennetto L., Visual attention preference for intermediate predictability in young children. Child Dev. 92, 691–703 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.Wu S., et al. , Macaques preferentially attend to intermediately surprising information. Biol. Lett. 18, 20220144 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.Gureckis T. M., Markant D. B., Self-directed learning: A cognitive and computational perspective. Perspect. Psychol. Sci. 7, 464–481 (2012). [DOI] [PubMed] [Google Scholar]
  • 15.A. Geana, R. Wilson, N. D. Daw, J. Cohen, “Boredom, information-seeking and exploration” in Proceedings of the Annual Meeting of the Cognitive Science Society (2016), vol. 38.
  • 16.Ten A., Kaushik P., Oudeyer P. Y., Gottlieb J., Humans monitor learning progress in curiosity-driven exploration. Nat. Commun. 12, 1–10 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Loewenstein G., The psychology of curiosity: A review and reinterpretation. Psychol. Bull. 116, 75 (1994). [Google Scholar]
  • 18.J. Schmidhuber, “Curious model-building control systems” in Proceedings of the International Joint Conference on Neural Networks (1991), pp. 1458–1463.
  • 19.Kidd C., Hayden B. Y., The psychology and neuroscience of curiosity. Neuron 88, 449–460 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.Dubey R., Griffiths T. L., Reconciling novelty and complexity through a rational analysis of curiosity. Psychol. Rev. 127, 455 (2020). [DOI] [PubMed] [Google Scholar]
  • 21.A. Geana, R. C. Wilson, N. Daw, J. D. Cohen, “Information-seeking, learning and the marginal value theorem: A normative approach to adaptive exploration” in Proceedings of the Annual Meeting of the Cognitive Science Society (2016), vol. 38.
  • 22.Csikszentmihalyi M., Flow: The Psychology of Optimal Experience (Harper & Row, New York, NY, 1990), vol. 1990. [Google Scholar]
  • 23.Melnikoff D. E., Carlson R. W., Stillman P. E., A computational theory of the subjective experience of flow. Nat. Commun. 13, 1–13 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24.Wilson R. C., Shenhav A., Straccia M., Cohen J. D., The eighty five percent rule for optimal learning. Nat. Commun. 10, 1–9 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25.J. A. Masís, D. Melnikoff, L. F. Barrett, J. Cohen, “When to choose: Information seeking in the speed-accuracy tradeoff” in Proceedings of the Annual Meeting of the Cognitive Science Society (2023), vol. 45.
  • 26.Agrawal M., Mattar M. G., Cohen J. D., Daw N. D., The temporal dynamics of opportunity costs: A normative account of cognitive fatigue and boredom. Psychol. Rev. 129, 564–585 (2021). [DOI] [PubMed] [Google Scholar]
  • 27.Y. Li, R. Carrasco-Davis, Y. Strittmatter, S. Sarao Mannelli, S. Musslick, “A meta-learning framework for rationalizing cognitive fatigue in neural systems” in Proceedings of the Annual Meeting of the Cognitive Science Society (2024), vol. 46.
  • 28.J. Masís, S. Musslick, J. D. Cohen, “The value of learning and cognitive control allocation” in Proceedings of the Annual Meeting of the Cognitive Science Society (2021), vol. 43.
  • 29.Ratcliff R., Rouder J. N., Modeling response times for two-choice decisions. Psychol. Sci. 9, 347–356 (1998). [Google Scholar]
  • 30.Ratcliff R., Smith P. L., Brown S. D., McKoon G., Diffusion decision model: Current issues and history. Trends Cogn. Sci. 20, 260–281 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31.Wiecki T. V., Sofer I., Frank M. J., HDDM: Hierarchical Bayesian estimation of the drift-diffusion model in Python. Front. Neuroinf. 7, 14 (2013). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32.Fengler A., Govindarajan L. N., Chen T., Frank M. J., Likelihood approximation networks (LANs) for fast inference of simulation models in cognitive neuroscience. eLife 10, e65074 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33.Shinn M., Lam N. H., Murray J. D., A flexible framework for simulating and fitting generalized drift-diffusion models. eLife 9, e56938 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34.Krajbich I., Lu D., Camerer C., Rangel A., The attentional drift-diffusion model extends to simple purchasing decisions. Front. Psychol. 3, 193 (2012). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 35.Callaway F., Rangel A., Griffiths T. L., Fixation patterns in simple choice reflect optimal information sampling. PLoS Comput. Biol. 17, e1008863 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 36.Cisek P., Puskas G. A., El-Murr S., Decisions in changing conditions: The urgency-gating model. J. Neurosci. 29, 11560–11571 (2009). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 37.Churchland A. K., Kiani R., Shadlen M. N., Decision-making with multiple alternatives. Nat. Neurosci. 11, 693–702 (2008). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 38.Drugowitsch J., Moreno-Bote R., Churchland A. K., Shadlen M. N., Pouget A., The cost of accumulating evidence in perceptual decision making. J. Neurosci. 32, 3612–3628 (2012). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 39.Hawkins G. E., Forstmann B. U., Wagenmakers E. J., Ratcliff R., Brown S. D., Revisiting the evidence for collapsing boundaries and urgency signals in perceptual decision-making. J. Neurosci. 35, 2476–2484 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 40.Simen P., Cohen J. D., Holmes P., Rapid decision threshold modulation by reward rate in a neural network. Neural Networks 19, 1013–1026 (2006). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 41.Bogacz R., Brown E., Moehlis J., Holmes P., Cohen J. D., The physics of optimal decision making: A formal analysis of models of performance in two-alternative forced-choice tasks. Psychol. Rev. 113, 700 (2006). [DOI] [PubMed] [Google Scholar]
  • 42.Newsome W. T., Pare E. B., A selective impairment of motion perception following lesions of the middle temporal visual area (MT). J. Neurosci. 8, 2201–2211 (1988). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 43.Newsome W. T., Britten K. H., Movshon J. A., Neuronal correlates of a perceptual decision. Nature 341, 52–54 (1989). [DOI] [PubMed] [Google Scholar]
  • 44.Shadlen M. N., Newsome W. T., Neural basis of a perceptual decision in the parietal cortex (area lip) of the rhesus monkey. J. Neurophysiol. 86, 1916–1936 (2001). [DOI] [PubMed] [Google Scholar]
  • 45.H. Ritz, J. DeGutis, M. J. Frank, M. Esterman, A. Shenhav, “An evidence accumulation model of motivational and developmental influences over sustained attention” in Proceedings of the Annual Meeting of the Cognitive Science Society (2020).
  • 46.Bieleke M., Barton L., Wolff W., Trajectories of boredom in self-control demanding tasks. Cogn. Emot. 35, 1018–1028 (2021). [DOI] [PubMed] [Google Scholar]
  • 47.Lin H., Saunders B., Friese M., Evans N. J., Inzlicht M., Strong effort manipulations reduce response caution: A preregistered reinvention of the ego-depletion paradigm. Psychol. Sci. 31, 531–547 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 48.R. Carrasco-Davis, J. Masís, A. M. Saxe, Meta-learning strategies through value maximization in neural networks. arXiv [Preprint] (2023). https://arxiv.org/abs/2310.19919 (Accessed 1 August 2024).
  • 49.Wilson R. C., Geana A., White J. M., Ludvig E. A., Cohen J. D., Humans use directed and random exploration to solve the explore-exploit dilemma. J. Exp. Psychol. Gen. 143, 2074 (2014). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 50.E. Schulz, E. D. Klenske, N. R. Bramley, M. Speekenbrink, “Strategic exploration in human adaptive control” in Proceedings of the Annual Meeting of the Cognitive Science Society (2017), vol. 39.
  • 51.D. Melnikoff, P. E. Stillman, R. W. Carlson, Optimal task representations for engagement, enjoyment, and performance. https://osf.io/preprints/osf/v93cp_v1. Accessed 1 August 2024.
  • 52.Lewis R. L., Howes A., Singh S., Computational rationality: Linking mechanism and behavior through bounded utility maximization. Top. Cogn. Sci. 6, 279–311 (2014). [DOI] [PubMed] [Google Scholar]
  • 53.Lieder F., Griffiths T. L., Resource-rational analysis: Understanding human cognition as the optimal use of limited computational resources. Behav. Brain Sci. 43, e1 (2020). [DOI] [PubMed] [Google Scholar]
  • 54.Tenenbaum J. B., Griffiths T. L., Generalization, similarity, and Bayesian inference. Behav. Brain Sci. 24, 629–640 (2001). [DOI] [PubMed] [Google Scholar]
  • 55.Körding K. P., Wolpert D. M., Bayesian integration in sensorimotor learning. Nature 427, 244–247 (2004). [DOI] [PubMed] [Google Scholar]
  • 56.Körding K. P., Wolpert D. M., Bayesian decision theory in sensorimotor control. Trends Cogn. Sci. 10, 319–326 (2006). [DOI] [PubMed] [Google Scholar]
  • 57.T. L. Griffiths, C. Kemp, J. B. Tenenbaum, “Bayesian models of cognition” in Cambridge Handbooks in Psychology, R. Sun, Ed. (Cambridge University Press, 2008), pp. 59–100.
  • 58.S. Ravi, S. Musslick, M. Hamin, T. Willke, J. D. Cohen, Navigating the tradeoff between multi-task learning and learning to multitask in deep neural networks. arXiv [Preprint] (2020). 10.48550/arXiv.2007.10527 (Accessed 1 August 2024). [DOI]
  • 59.Lieder F., Shenhav A., Musslick S., Griffiths T. L., Rational metareasoning and the plasticity of cognitive control. PLoS Comput. Biol. 14, e1006043 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 60.Bustamante L., Lieder F., Musslick S., Shenhav A., Cohen J., Learning to overexert cognitive control in a stroop task. Cogn., Affect., Behav. Neurosci. 21, 453–471 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 61.Giallanza T., Campbell D., Cohen J. D., Toward the emergence of intelligent control: Episodic generalization and optimization. Open Mind 8, 688–722 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 62.Leng X., Yee D., Ritz H., Shenhav A., Dissociable influences of reward and punishment on adaptive cognitive control. PLoS Comput. Biol. 17, e1009737 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 63.Gentile C., Warmuth M. K., Linear hinge loss and average margin. Adv. Neural Inf. Process. Syst. 11, 225–231 (1998). [Google Scholar]
  • 64.Zacksenhouse M., Bogacz R., Holmes P., Robust versus optimal strategies for two-alternative forced choice tasks. J. Math. Psychol. 54, 230–246 (2010). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 65.Bogacz R., Hu P. T., Holmes P. J., Cohen J. D., Do humans produce the speed-accuracy trade-off that maximizes reward rate? Q. J. Exp. Psychol. 63, 863–891 (2010). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 66.Balci F., et al. , Acquisition of decision making criteria: Reward rate ultimately beats accuracy. Atten., Percept., Psychophys. 73, 640–657 (2011). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 67.Gold J. I., Shadlen M. N., Banburismus and the brain: Decoding the relationship between sensory stimuli, decisions, and reward. Neuron 36, 299–308 (2002). [DOI] [PubMed] [Google Scholar]
  • 68.H. Ritz, R. Frömer, A. Shenhav, Phantom controllers: Misspecified models create the false appearance of adaptive control during value-based choice. bioRxiv [Preprint] (2023). 10.1101/2023.01.18.524640 (Accessed 1 August 2024). [DOI]
  • 69.J. A. Masis, S. Musslick, J. D. Cohen, Data from “Learning expectations shape cognitive control allocation." Open Science Framework. https://osf.io/w46ev. Deposited 4 April 2025. [DOI] [PMC free article] [PubMed]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Appendix 01 (PDF)

pnas.2416720122.sapp.pdf (56.1MB, pdf)

Data Availability Statement

Data and code are available at https://osf.io/w46ev/ (69). A summary of the software used can be found in SI Appendix.


Articles from Proceedings of the National Academy of Sciences of the United States of America are provided here courtesy of National Academy of Sciences

RESOURCES