Abstract
Learning rules—prescriptions for updating model parameters to improve performance—are typically assumed rather than derived. Why do some learning rules work better than others, and under what assumptions can a given rule be considered optimal? We propose a theoretical framework that casts learning rules as policies for navigating (partially observable) loss landscapes, and identifies optimal rules as solutions to an associated optimal control problem. A range of well-known rules emerge naturally within this framework under different assumptions: gradient descent from short-horizon optimization, momentum from longer-horizon planning, natural gradients from accounting for parameter space geometry, non-gradient rules from partial controllability, and adaptive optimizers like Adam from online Bayesian inference of loss landscape shape. We further show that continual learning strategies like weight resetting can be understood as optimal responses to task uncertainty. By unifying these phenomena under a single objective, our framework clarifies the computational structure of learning and offers a principled foundation for designing adaptive algorithms.
1. Introduction
A central concern in machine learning is identifying parameters that optimize model performance. Because directly searching for optimal parameters (e.g., via a grid search) is prohibitively costly in the high-dimensional parameter spaces characteristic of models based on artificial neural networks, optimization typically involves seeking iterative improvements in performance rather than directly searching for a global optimum. Procedures for iterative parameter improvement, or learning rules, are most commonly some variant of gradient descent [1, 2], with the backpropagation algorithm [3, 4] being a notable example. In biological neural networks, the plausibility of gradient descent is hotly debated [5, 6], and alternative rules that do not follow gradients have been proposed [7, 8].
Variants of gradient descent can largely be classified in terms of the presence or absence of three elements: momentum [9, 10], an adaptive learning rate [11-13], and loss approximation. From a geometric perspective, gradient descent corresponds to moving down the steepest part of the local loss landscape. In this view, momentum helps ensure smooth parameter changes, even when the loss landscape changes abruptly; an adaptive learning rate allows parameters to change more quickly when gradients are steady, or equivalently in regions where the loss landscape is flat; and computing the loss or its gradients approximately, for example over mini-batches, both improves efficiency and may add noise useful for generalization [14-16].
Why prefer one learning rule over another? Is a variant of gradient descent always optimal? If so, which one? If not, by what criteria do we construct or decide on an alternative? These questions are usually answered empirically, with optimizers like Adam [13] popular because they have proven performant in a wide variety of contexts [17-20]. It would be helpful to have a normative framework for answering them in a principled fashion, which, given a set of assumptions, identifies some learning rule as ‘optimal’. In this paper, our aim is to provide such a unifying framework.
Our three key insights are as follows. First, one can improve performance by optimizing not just over the next small parameter update, but over the whole sequence of future updates; this allows optimization to be less ‘myopic’ and more ‘farsighted’. Second, optimal learning dynamics ought to be sensitive to structure in parameter space. This insight is related to, but slightly more general than, the line of thought that leads to natural gradient descent [21]. Third, one can view the loss landscape as being only partially observable, which implies optimal learning ought to depend on beliefs about the loss landscape. Assuming partial observability is a useful way to model the fact that the training loss is typically only a proxy for the test loss, which is the true optimization target. We show this idea naturally yields adaptive optimizers like Adam, which update parameters using inferred loss shape.
Our framework is significant for two reasons. First, it makes it easier to generate new learning rules from a principled starting point, and to justify them without empirical guesswork. Second, it helps clarify which features of existing learning rules are essential for performance, and which are incidental. In the following sections, we discuss in more detail how different classes of well-known learning rules—including gradient descent with momentum (Sec. 3), natural gradient descent (Sec. 4), and rules with adaptive learning rates (Sec. 5)—can all be derived from our framework. Finally, Sec. 6 uses our framework to justify recently-identified rules for continual learning [22].
2. Mathematical formulation: learning rules as loss landscape navigation
Gradient descent and Newton’s method as optimal single steps through parameter space.
To motivate our framework, it is useful to observe that gradient descent and Newton’s method (a second-order analogue [23]) minimize a certain objective. Given a loss that depends on a parameter vector , and a local first- or second-order approximation of it, we would like an update that decreases the loss as much as possible. Since large enough updates would invalidate our local loss approximation, we want to do this subject to the constraint that is not too large. We can do this by minimizing a combination of the loss and a step-size-related regularization term:
where is the learning rate (which determines the size of the ‘trust region’), and is the Hessian of at . Given the first-order loss approximation, the solution is , i.e., gradient descent with a learning rate . Geometrically, this step is down the ‘steepest’ part of the loss landscape near . Given the second-order loss approximation, the solution is , i.e., gradient descent with a curvature-sensitive effective learning rate . One still takes a step down the steepest part of the loss landscape, but takes a larger step if the landscape is very flat; this avoids the slowdown gradient descent faces near a local minimum, where gradients get smaller. Newton’s method technically only corresponds to the large limit, where ; more generally, we obtain a regularized version, which (given a specific Hessian approximation) is sometimes called the Levenberg-Marquardt algorithm [24].
Learning rules as partially observable loss landscape navigation.
The single-step view of loss optimization is arguably myopic, since the best short-term improvement of the loss may be suboptimal in the long term; famously, gradient descent finds local rather than global minima. An obvious way to address this issue is to instead optimize over multiple steps, and ask what sequence of steps should be taken in order to minimize the loss (Fig. 1a, b). In going from single- to multi-step optimization, we convert the problem of deciding on parameter updates into a navigation problem: what path through the loss landscape minimizes the loss, while simultaneously avoiding overly fast parameter changes?
Figure 1: Basic idea of our framework and a simple example.

a. Single-step approaches (top) optimize over short-term changes to the loss, while a multi-step approach (bottom) optimizes over longer-term changes. b. Gradient descent vs multi-step optimization for a double-well loss, with the optimal multi-step trajectory computed by directly minimizing the objective (see Appendix B for details). Note that the multi-step rule converges to the global rather than local minimum. c. Values of the kinetic (top) and potential (bottom) terms along the optimal trajectory from (b). The loss/potential does not decrease monotonically, since the learner must first escape a local minimum.
To formalize this, we define a continuous-time optimal control problem [25] over learning trajectories:
| (1) |
where is the learning rate, weights the influence of the loss, and is the temporal discounting rate. This objective generalizes the single-step objective (see Appendix A) and defines a continuous-time reinforcement learning [26, 27] and control [25] problem. It penalizes a combination of abrupt parameter changes (through the first term) and high loss (through the second), while discounting temporally distant costs through the factor . This objective also effectively turns learning into a physics problem [28], with these two terms analogous to kinetic and potential energy, and (the parameter space metric) and (the ‘drift’ or ‘bias’) encoding assumptions about the ambient (parameter space) geometry. Note: the potential/loss term may not decrease monotonically, for example if optimal dynamics involves exiting a local minimum of the loss (Fig. 1c).
Why partial observability?
In most practical settings, the learner does not have full knowledge of the loss landscape. We model this by assuming the learner possesses a structured model of that can evolve in time, which is informed by past observations (e.g., of gradients and curvature). This partial observability assumption is realistic for three reasons. First, due to the sheer size of parameter space, it is difficult to plan using more than a small part of the loss landscape at any given time. Second, the training loss is generally different from the test loss, but the two are not unrelated; it can be helpful to view the training loss as a noise-corrupted version of the test loss. Finally, the loss is usually evaluated on a batch of examples rather than the full training set. While these constraints make it difficult to plan an entire trajectory in one shot, one can instead work iteratively: take a step in the optimal direction, re-estimate your belief given new observations, and repeat.
Deriving learning rules.
Our objective (Eq. 1) takes as input four pieces of data: a loss function , a parameter space metric , a drift term , and a random variable , which models the learner’s belief about the loss landscape. This belief can evolve in time, e.g., via online Bayesian inference. Given this data, an optimal learning trajectory is a choice of that minimizes . Given an initial state and a desired time step , an optimal learning rule is the difference along the optimal trajectory, which is proportional to the initial velocity when is small. This defines a navigation policy and continuous-time analogue of the optimal ‘next step’. The naive way to estimate an optimal trajectory is to consider a parameterized family of trajectories, and then directly minimize with respect to those parameters; Strang et al. [29] use this type of approach to numerically solve classical mechanics problems, and we followed their approach to generate Fig. 1b, c (see Appendix B for details). But there is a powerful alternative approach: framing learning rule optimization in terms of continuous-time learning trajectories allows us to leverage a result from the calculus of variations [28, 30] to identify optimal dynamics with solutions of the Euler-Lagrange (EL) equations (see Appendix C). These second-order ODEs are often solvable, and in our case, yield familiar learning rules like gradient descent and Adam under different assumptions. See Appendix D for discussion of technical and conceptual subtleties related to boundary conditions.
3. Momentum as a generic consequence of multi-step trajectory optimization
Unlike gradient descent, learning rules with momentum [9, 31] are more ‘inertial’: instead of following the current gradient, one follows a weighted combination of current and past gradients. Like second-order methods, this change helps speed up movement through flat regions of the loss landscape. In this section, we show that momentum is a generic consequence of multi-step optimization, and derive first- and second-order learning rules with momentum using our framework.
Momentum as a generic consequence of multi-step optimization.
A straightforward, multi-step generalization of the objective we considered to justify gradient descent is (see Appendix E)
| (2) |
This objective is the simplest version of Eq. 1 (, , and the loss is not approximated). It penalizes only two things: abrupt parameter changes, and the loss. Optimal learning dynamics follow the EL equations, which in this case take the simple form (see Appendix E)
| (3) |
where we defined2 momentum as . These equations have two interesting consequences. First, we obtain momentum essentially for free, simply from going from one-step to multi-step optimization. Second, the temporal discounting rate allows one to interpolate between a momentum-based rule and standard gradient descent, since in the ‘overdamped’ (large ) limit these equations become
| (4) |
This objective has a straightforward analogy with a mechanical system [28]: a particle with a mass moves in a potential and experiences an amount of friction proportional to . When ‘friction’ is sufficiently high, dynamics become non-inertial, and dominated by the shape of the potential/loss.
Deriving first- and second-order learning rules with and without momentum.
We can derive first- and second-order learning rules by tweaking Eq. 2 to locally approximate near as
| (5) |
where and are the gradient and Hessian of at . For any , we find (Appendix E) that
| (6) |
is the optimal update. See Fig. 2a for example traces and Fig. 2b for example loss traces. Importantly, we can derive three learning rules from this result. When the step size is sufficiently large, we recover the Hessian-conditioned rule (i.e., Newton’s method) that ‘jumps’ straight to the minimum at . When is somewhat smaller than the other characteristic time scales, we either get gradient descent (in the overdamped large limit) or a ballistic learning rule:
| (7) |
Figure 2: Effect of modulating temporal discounting rate.

a. Example optimal traces for a 1D quadratic loss, assuming different values of the temporal discounting rate (, 1, 10). b. Loss over time given the from (a), same values of . Note that lower values of (‘longer’ planning horizon) produce loss curves that converge more quickly. c. Shape of trajectories for different given a 2D anisotropic loss. In the gradient-descent-like regime (), converges much more quickly than due to the anisotropy. In the ballistic () regime, the difference in convergence rates is not as extreme. d. Ratio of convergence rates assuming a diagonal Hessian. In the gradient-descent-like regime, directions with four times as much curvature converge 4 times faster; in the ballistic regime, they only converge times faster. e. An Adam-like implementation of the ballistic () rule was used to train a small multilayer perceptron (MLP) to classify MNIST digits. Left: loss over training, right: test set accuracy over training. f. Same as in (e), but for a small convolutional neural network (CNN) trained to classify CIFAR-10 images. The ballistic rule generally performs better than SGD (black), and similarly to or worse than Adam (red).
The last learning rule is not usually considered, but has a a square-root normalization similar to Adam [13]. We call such a rule ballistic because it corresponds to the frictionless limit; it strikes a compromise between following the gradient, as gradient descent does, and jumping straight to the local minimum of the loss, as Newton’s method does. Because it depends on the square root of the Hessian, differences in how quickly different parameter directions converge are reduced somewhat relative to gradient descent (Fig. 2c, d). A heuristic implementation of this rule performs well on standard datasets (Fig. 2e, f), which suggests its behavior may be reasonable even for non-quadratic losses. See Appendix B for experiment details.
4. Parameter space geometry modulates optimal learning dynamics
Parameter space geometry modulates distances.
An insight due to Amari [21] is that learning rules ought to be sensitive to the structure of parameter space. If distances in parameter space follow a non-Euclidean metric (e.g., the Fisher information matrix), the objective we used to derive gradient descent must be modified to involve a parameter change penalty , which causes the optimal first-order (one-step) rule to become . We can implement Amari’s insight in the continuous-time, multi-step setting by adding a metric to Eq. 2:
| (8) |
The corresponding EL equations that describe optimal learning dynamics read (see Appendix F)
| (9) |
and by an argument analogous to the one in the previous section, we recover three types of rules:
| (10) |
These follow from making the replacements and in a model that assumes is locally -independent (i.e., here, refers to ). In addition to recovering natural gradient descent (middle), which modifies gradient descent to account for non-Euclidean parameter space geometry [21], we also recover two other learning rules. The first is Newton’s method with a metric- and Hessian-dependent learning rate, and the last is a ballistic method which looks like Newton’s method, but involves instead of . This method compromises between using and as preconditioners, or equivalently between moving based on parameter space and loss landscape geometry. While prior work has attempted to combine natural gradients and momentum heuristically [32], our framework shows how these elements arise jointly from a principled objective. As is well-known, using an anisotropic can add anisotropy to learning dynamics (Fig. 3a, b).
Figure 3: Parameter space geometry affects optimal learning trajectories.

a. Optimal trajectory through space for an isotropic quadratic loss, assuming no nontrivial and . The heatmap and contours show the value of the loss at each (, ) value. Black line: optimal trajectory, red dot: global minimum of loss. Note that, because the loss is isotropic, the optimal trajectory is too. b. Same as (a), but given a strongly anisotropic constant metric . Note that the optimal trajectory is no longer the same along each direction, but converges much more quickly along the direction. c. Same as (a), but given that corresponds to purely rotational dynamics. Note two differences: it spirals about the origin, and no longer converges to the global minimum of the loss, but to a different point closer to the origin (orange dot). d. Same as (a), but given that corresponds to weight decay. There is no anisotropy, but the trajectory does not converge to the minimum of the loss.
Natural gradient descent is not second-order optimization in disguise.
It is often argued that natural gradient descent behaves like a second-order method, with the metric acting as a surrogate for the Hessian . Martens [33] explores this view in detail, and shows that natural gradient descent is sometimes equivalent to a Generalized Gauss-Newton method. But Martens also notes a variety of problems with this view, like the fact that existing theory (e.g., convergence rates) portrays the Hessian as more performant, while empirical work shows natural gradient descent is more performant. Our results suggest that the analogy between second-order methods and natural gradient descent is misguided, and that and play fundamentally different roles: governs the parameter velocity penalty, while measures loss landscape curvature. Physically, the former defines ambient geometry (as in general relativity [34]), whereas the latter pushes and pulls particles along that geometry. It is also clear when one examines the different roles of and in the rules we derived above.
Optimizing partially controllable parameters.
Parameter space geometry can also influence what kind of learning rule is optimal in a different, less well-appreciated way: suppose an external ‘force’ pushes on the parameter in a state-dependent fashion, like a ‘wind’ that determines which directions are more or less difficult to travel in. This feature can arise from partial controllability; if our control of a parameter is partial, e.g., , controlling the parameter via involves not just moving in a direction that improves the loss, but also fighting against the ‘default’ dynamics . Here, we show that this feature can produce non-gradient rules. If we add a drift term to Eq. 8,
| (11) |
The presence of and can greatly complexify the EL equations, especially if they are state-dependent. Consider the effect of the drift by itself (i.e., assume ; see Appendix G):
where is the Jacobian of at . One can interpret as contributing to dynamics in two ways: it contributes an effective ‘potential-like’ term , or equivalently an effective loss; and it produces an effective state-dependent discounting factor . Note that this reduces to the usual discounting factor if for all , which is true if and only if is the gradient of some function.
In Appendix G, we consider two examples in more detail: the case where corresponds to rotational dynamics (Fig. 3c), and the case where corresponds to weight decay (Fig. 3d). The former case is ‘nonconservative’ (e.g., ) but the second is not. Interestingly, rotational dynamics can make ‘spiraling’ trajectories optimal, and both choices of make optimal trajectories converge to a point different from the global minimum of the loss.
Optimal learning dynamics are generally non-gradient.
One significant consequence of including a drift term is that it can cause the optimal learning rule to involve non-gradient dynamics, which cannot be described as gradient descent (with or without momentum) down any objective. In our example, this happens if and only if itself is not the gradient of any function (or equivalently, if is not symmetric for all ). In terms of a Helmholtz decomposition [35] , where is some non-unique ‘potential’ function and is divergence-free (i.e., ), is the interesting component of . See Appendix G for more discussion of this point. This interpretation suggests a concrete diagnostic: when learning rules involve update components that do not point downhill, such as decay terms or rotational drift, they may be optimal responses to implicit background dynamics.
5. Adaptive optimizers arise from beliefs about how gradients evolve in time
Adaptive optimizers like Adam [13] and RMSprop [12] use a learning rate that depends on the variance of recent gradients. If recent gradients were consistent, learn quickly; if not, make smaller updates. Doing this is known to work well in practice, with Adam variants outperforming most other known optimizers [17-20] in typical settings, but it is unclear why, or under what assumptions Adam is optimal. In this section, we show that an Adam-like strategy is optimal under a certain Bayesian model of how the local loss landscape shape evolves in time. More generally, we show adaptive optimizers arise from using past observations of landscape shape to estimate current and future shape.
Modeling uncertainty in the time evolution of the local loss landscape.
One strategy for navigating a partially observable loss landscape is to maintain running estimates of quantities that characterize its local shape, like the gradient and curvature. Motivated by this idea, assume that the learner maintains a local model of loss landscape shape near their current location , i.e.,
| (12) |
where the gradient estimate and Hessian estimate are assumed to be time-varying. Here, is a fixed (known) scaling factor. Since and represent average values of gradients and curvature, respectively, they are not directly observable—but the learner can estimate them via their assumed link to observable gradients . Assume that the learner has a Bayesian model of how local landscape geometry evolves with two components: an observation model, which connects to and ; and a prior belief about how landscape geometry evolves.
Motivated by a simple model of gradient drift and diffusion in a quadratic loss (see Appendix H), assume that , which implies that and . To make inference simpler, we will use a crude approximation3 associated with method-of-moment-based strategies, and assume and provide independent observations of and , i.e., and . Here, and are ‘observation noise’ parameters.
One reasonable prior, which assumes that both parameters tend towards zero and become more uncertain in the absence of observations, is an Ornstein-Uhlenbeck prior. Assume
| (13) |
where and parameterize decay, and parameterize the uncertainty growth rate, and and are Gaussian white noise terms. This means and .
Time-evolving beliefs yield adaptive optimal learning rule.
Assume an objective that prioritizes a mix of small loss, smooth parameter changes, and good inference (implemented via a term):
where is the learner’s posterior belief about local landscape shape dynamics. This objective has a well-defined limit (see Appendix H) with EL equations
These equations are somewhat complicated, but considerably simplify if one makes assumptions about the relative sizes of parameters. If we assume that parameter changes are ballistic () but landscape beliefs change somewhat more slowly (, ),
To good approximation, this means that the optimal learning rule has
| (14) |
Note that the effective learning rate goes like rather than . Our framework identifies the square root as a consequence of assuming highly inertial parameter changes, but gradient-descentlike landscape shape estimate dynamics. Moreover, contrary to other ideas about Adam [36], the square root is a feature rather than a bug, and need not be ‘fixed’. Other features of Adam, like estimating using the uncentered averages of and approximating as diagonal, appear to be approximations that improve efficiency and scalability, as is usually believed.
6. Noisy continual learning strategies reflect parsimony and task uncertainty
In continual learning settings, agents must balance their ability to learn new tasks with their ability to retain information about previous tasks [37, 38]. Poorly performing agents exhibit catastrophic forgetting, but even in the absence of such forgetting, many algorithms exhibit a progressive loss in plasticity as an ever larger number of tasks are learned [22]. Dohare et al. observed that periodically resetting weights that do not strongly respond to task gradients empirically seems to ameliorate this issue. Others have observed that injecting noise in different ways, like via randomly perturbing little-used weights, also helps address this issue [39]. Thus far, it has been somewhat unclear why any of these strategies ought to work, and to what extent various details (e.g., how noise is injected) matter. Our framework provides qualitative guidance here: the reason different noise injection strategies empirically work is that injecting noise at all is more important than how that noise is injected.
Modeling weight-uncertainty-sensitive learning dynamics.
Assume that the learner uses distributional estimates, rather than point estimates, of the model parameters . For simplicity, assume each is associated with a normal distribution , where denotes the mean estimate of , and denotes the posterior variance. Instead of penalizing abrupt weight changes, in this setting it makes more sense to penalize abrupt changes in weight distribution:
where we have also included an entropy term to explicitly penalize ‘model complexity’. When written more explicitly, this objective reduces to our standard form (Eq. 1) with a nontrivial metric (see Appendix I). Given a local (quadratic) loss approximation with gradient and Hessian , the EL equations read
| (15) |
| (16) |
The equation for (Eq. 15) has an interesting feature: it involves a variance-dependent effective discounting rate and effective learning rate . One can interpret the former as saying that planning should become more short-term when weight variances are changing quickly (i.e., is high), and that learning should speed up when uncertainty (i.e., ) is high.
Variance reflects both parsimony and task-driven instability.
The equation for (Eq. 16) specifies how variance ought to increase or decrease along the optimal learning trajectory. The two terms on the right-hand side tell us how this happens: variance decreases if the nearby loss landscape is very curved (intuitively, the nearest minimum is easier to find); the entropic term incentivizes increases in variance, in order to favor the ‘simplest’ model associated with a given loss value; and the final term affects variance in a more subtle way. It increases variance if the mean is changing more quickly than the variance, something that can happen in continual learning settings when there is a transition between tasks. If the variance is changing quickly but the mean is not changing, it decreases variance somewhat. See Appendix I for more discussion.
The entropic term enforces behavior analogous to the weight resetting of Dohare et al. [22]: when the loss is not particularly sensitive to a given weight’s value (i.e., is small), uncertainty about that weight should increase, with one possible mechanism being a reset. The last term on the right-hand side, which compares the speed of mean and variance changes, appears sensitive to task uncertainty. Together, these two effects—one tied to parsimony, the other to volatility—help explain the behavior of continual learning algorithms that inject noise or reset unused weights, like Dohare et al.’s method.
7. Discussion
We proposed a unifying framework that treats learning rules as solutions to an optimal control problem. By varying assumptions about geometry, planning horizon, and uncertainty, it recovers a wide range of familiar rules—including gradient descent, momentum, natural gradient descent, Adam, and continual learning strategies—from a single objective. This framework rests on three core ideas: learning unfolds over multiple steps, not one; parameter space can have nontrivial geometry and dynamics; and the loss landscape may only be partially observable, in which case it must be inferred. Our framework not only recovers these components individually, but also shows how they can be combined, yielding principled algorithms that integrate elements like momentum and natural gradients. It also suggests a shift in how learning rules should be evaluated: rather than focusing only on convergence rates or optimization guarantees, we should consider what assumptions a rule implicitly encodes and what problem it is actually solving (e.g., one related to generalization). This perspective is especially relevant for adaptive optimizers like Adam, which our framework portrays as optimizing an inferred test loss based on a specific Bayesian model of loss landscape shape dynamics. In many deep learning settings, performant optimizers ought to get things slightly ‘wrong’ to generalize [14, 40-42].
Connection to physics.
Our optimal control formulation of learning dynamics can be precisely mapped to a classical mechanics problem, which may allow physics tools to be adapted to study learning algorithms. For example, Noether’s theorem [43] implies that (quasi-) symmetries of objectives like Eq. 1 imply conserved quantities along optimal trajectories. Analogous ideas linking the consequences of symmetry to noisy recurrent dynamics [44], learning dynamics [45], and optimal dimensionality reduction [46] have already begun to be explored.
Biological relevance.
Neural circuits in the brain are thought to implement learning rules—such as Hebbian plasticity, homeostatic mechanisms, or spike-timing-dependent plasticity—that often lack a clear gradient-descent interpretation [7, 8, 47]. These rules may include time-asymmetric or rotational dynamics, operate under metabolic or architectural constraints, or adjust weights in response to activity thresholds rather than loss functions. Our framework suggests that such rules may still be optimal under constraints imposed by biological dynamics, such as synaptic decay, intrinsic drift, or limited access to global error signals [5, 48].
By treating learning as constrained control in a partially observable environment, our framework offers a normative lens on how non-gradient rules can potentially emerge as efficient strategies in biologically realistic regimes. This perspective is analogous to and consistent with observations that complex behavioral strategies can emerge from a mix of simple rewards and naturalistic constraints (e.g., partial observability), for example in foraging tasks [49].
Related work.
Several recent efforts aimed to unify learning rules under broader frameworks. Khan and Rue [50] portray various optimizers as natural gradient descent with respect to a fairly general objective, and Shoji et al. [51] similarly observe that many learning rules can be viewed as instances of natural gradient descent. However, both works take the idea that natural gradient descent is optimal for granted, and do not incorporate multi-step planning, belief updating, or partial observability.
Our framework has parallels with earlier foundational work by Wibisono et al. [52] which relates accelerated optimization methods to a continuous-time variational objective. However, there are also important qualitative differences between their framework and ours, with the most important being that their objective cannot be interpreted as a sum of costs, unlike ours. Our objective contains two terms—a quadratic parameter velocity penalty (or ‘kinetic energy’), and a loss term (or ‘potential energy’)—which are added together. In their objective, as in classical mechanics, the potential term is subtracted from the kinetic term. While this sign difference means that their approach is more directly related to classical mechanics, it also means that it is more distantly related to optimal control. A practical benefit of our sign choice is that optimal trajectories minimize the objective, rather than make it stationary in general.
Concurrent work by Orvieto and Gower [53] also proposes a view of Adam related to loss landscape shape inference, and also formalizes it in terms of an objective with a term, but our picture and theirs differ in certain details. Perhaps the most important is that we link the appearance of the square root of the Hessian to operating in the ‘ballistic’ regime, or equivalently longer-term planning.
Our work is related in spirit to recent efforts to unify training objectives in deep learning, such as the framework proposed by Alshammari et al. [54].
Limitations.
We consider a setting in which our notion of ‘optimal’ does not factor in concerns which often affect optimization in practice, like memory requirements, computational simplicity, and efficiency. Relatedly, even if a rule is identified as optimal given our framework, it is unclear how useful it may be in practice, especially if it requires potentially expensive matrix operations. On the other hand, given that our formulation is in terms of an objective function, it may be possible to penalize things like efficiency explicitly in order to partially address this issue.
Finally, we do not claim to be able to explain all phenomena related to learning dynamics or learning rules; our goal in this work is merely to propose a useful framework for thinking about why features of well-known learning rules (like momentum) might be useful. The most important consequence of our framework is that, given an assumed objective (e.g., with a certain amount of temporal discounting, and a particular model of landscape shape belief updating), any two learning rules can be compared, and one or more learning rules can be shown to be optimal. Whether and how to link this framework to specific empirical circumstances is a different question which we expect to be more difficult.
Supplementary Material
Acknowledgments and Disclosure of Funding
SJG was funded by the Kempner Institute for the Study of Natural and Artificial Intelligence, and a Polymath Award from Schmidt Sciences. KR was funded by the NIH (RF1DA056403, U01NS136507), James S. McDonnell Foundation (220020466), Simons Foundation (Pilot Extension-00003332-02), McKnight Endowment Fund, CIFAR Azrieli Global Scholar Program, NSF (2046583), a Harvard Medical School Neurobiology Lefler Small Grant Award, and a Harvard Medical School Dean’s Innovation Award. This work has been made possible in part by a gift from the Chan Zuckerberg Initiative Foundation to establish the Kempner Institute for the Study of Natural and Artificial Intelligence at Harvard University.
Footnotes
This convention matches machine learning practice. To better parallel physics, we would choose .
See Appendix H for a discussion of why this is reasonable, and for a description of the somewhat more complicated rule one gets if one does not assume this.
References
- [1].Curry Haskell B.. The method of steepest descent for non-linear minimization problems. Quarterly of Applied Mathematics, 2(3):258–261, 1944. [Google Scholar]
- [2].Ruder Sebastian. An overview of gradient descent optimization algorithms. arXiv e-prints, page arXiv:1609.04747, September 2016. [Google Scholar]
- [3].Rumelhart DE, Hinton GE, and Williams RJ. Learning internal representations by error propagation, page 318–362. MIT Press, Cambridge, MA, USA, 1986. [Google Scholar]
- [4].Rumelhart David E., Hinton Geoffrey E., and Williams Ronald J.. Learning representations by back-propagating errors. Nature, 323(6088):533–536, October 1986. [Google Scholar]
- [5].Lillicrap Timothy P., Cownden Daniel, Tweed Douglas B., and Akerman Colin J.. Random synaptic feedback weights support error backpropagation for deep learning. Nature Communications, 7(1):13276, November 2016. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [6].Lillicrap Timothy P., Santoro Adam, Marris Luke, Akerman Colin J., and Hinton Geoffrey. Backpropagation and the brain. Nature Reviews Neuroscience, 21(6):335–346, June 2020. [DOI] [PubMed] [Google Scholar]
- [7].Hebb Donald O.. The organization of behavior: A neuropsychological theory. New York: Wiley and Sons, 1949. [Google Scholar]
- [8].Gerstner Wulfram and Kistler Werner M.. Mathematical formulations of Hebbian learning. Biological Cybernetics, 87(5):404–415, December 2002. [DOI] [PubMed] [Google Scholar]
- [9].Polyak BT. Some methods of speeding up the convergence of iteration methods. USSR Computational Mathematics and Mathematical Physics, 4(5):1–17, 1964. [Google Scholar]
- [10].Sutskever Ilya, Martens James, Dahl George, and Hinton Geoffrey. On the importance of initialization and momentum in deep learning. In Dasgupta Sanjoy and McAllester David, editors, Proceedings of the 30th International Conference on Machine Learning, volume 28 of Proceedings of Machine Learning Research, pages 1139–1147, Atlanta, Georgia, USA, 17–19 Jun 2013. PMLR. [Google Scholar]
- [11].Duchi John, Hazan Elad, and Singer Yoram. Adaptive subgradient methods for online learning and stochastic optimization. Journal of Machine Learning Research, 12(61):2121–2159, 2011. [Google Scholar]
- [12].Tieleman T and Hinton G. Rmsprop: Divide the gradient by a running average of its recent magnitude. Coursera: Neural Networks for Machine Learning, 2012. [Google Scholar]
- [13].Kingma Diederik P. and Ba Jimmy. Adam: A method for stochastic optimization. In Third International Conference on Learning Representations, 2015. [Google Scholar]
- [14].Keskar Nitish Shirish, Mudigere Dheevatsa, Nocedal Jorge, Smelyanskiy Mikhail, and Tang Ping Tak Peter. On large-batch training for deep learning: Generalization gap and sharp minima. In International Conference on Learning Representations, 2017. [Google Scholar]
- [15].Smith Samuel L. and Le Quoc V.. A bayesian perspective on generalization and stochastic gradient descent. In International Conference on Learning Representations, 2018. [Google Scholar]
- [16].Smith Samuel, Elsen Erich, and De Soham. On the generalization benefit of noise in stochastic gradient descent. In Daumé Hal III and Singh Aarti, editors, Proceedings of the 37th International Conference on Machine Learning, volume 119 of Proceedings of Machine Learning Research, pages 9058–9067. PMLR, 13–18 Jul 2020. [Google Scholar]
- [17].Choi Dami, Shallue Christopher J., Nado Zachary, Lee Jaehoon, Maddison Chris J., and Dahl George E.. On Empirical Comparisons of Optimizers for Deep Learning. arXiv e-prints, page arXiv:1910.05446, October 2019. [Google Scholar]
- [18].Soydaner Derya. A comparison of optimization algorithms for deep learning. International Journal of Pattern Recognition and Artificial Intelligence, 34(13):2052013, 2020. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [19].Kunstner Frederik, Milligan Alan, Yadav Robin, Schmidt Mark, and Bietti Alberto. Heavytailed class imbalance and why adam outperforms gradient descent on language models. In Globerson A, Mackey L, Belgrave D, Fan A, Paquet U, Tomczak J, and Zhang C, editors, Advances in Neural Information Processing Systems, volume 37, pages 30106–30148. Curran Associates, Inc., 2024. [Google Scholar]
- [20].Morrison Nathaniel and Ma Eric Y.. Efficiency of machine learning optimizers and meta-optimization for nanophotonic inverse design tasks. APL Machine Learning, 3(1):016101, January 2025. _eprint: https://pubs.aip.org/aip/aml/article-pdf/doi/10.1063/5.0238444/20328279/016101_1_5.0238444.pdf. [Google Scholar]
- [21].Amari Shun-ichi. Natural gradient works efficiently in learning. Neural Computation, 10(2):251–276, 1998. [Google Scholar]
- [22].Dohare Shibhansh, Hernandez-Garcia J. Fernando, Lan Qingfeng, Rahman Parash, Mahmood A. Rupam, and Sutton Richard S.. Loss of plasticity in deep continual learning. Nature, 632(8026):768–774, August 2024. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [23].Deuflhard Peter. Newton Methods for Nonlinear Problems: Affine Invariance and Adaptive Algorithms. Springer Berlin Heidelberg, Berlin, Heidelberg, 2011. [Google Scholar]
- [24].Nocedal Stephen J. Wright Jorge. Numerical Optimization. Springer New York, New York, NY, 2006. [Google Scholar]
- [25].Bertsekas Dimitri. Dynamic programming and optimal control: Volume I. Athena scientific, 4th edition, 2012. [Google Scholar]
- [26].Sutton Richard S. and Barto Andrew G.. Reinforcement Learning: An Introduction. MIT Press, Cambridge, MA, 2nd edition, 2018. [Google Scholar]
- [27].Doya Kenji. Reinforcement Learning in Continuous Time and Space. Neural Computation, 12(1):219–245, January 2000. _eprint: https://direct.mit.edu/neco/article-pdf/12/1/219/814284/089976600300015961.pdf. [DOI] [PubMed] [Google Scholar]
- [28].Goldstein Herbert. Classical mechanics. Addison-Wesley, 1980. [Google Scholar]
- [29].Strang Tim, Caruso Isabella, and Greydanus Sam. Nature’s Cost Function: Simulating Physics by Minimizing the Action. arXiv e-prints, page arXiv:2303.02115, March 2023. [Google Scholar]
- [30].Elsgolc Lev D.. Calculus of variations. Dover, 2007. [Google Scholar]
- [31].Murphy Kevin P.. Probabilistic Machine Learning: An introduction. MIT Press, 2022. [Google Scholar]
- [32].Khan Mohammad, Nielsen Didrik, Tangkaratt Voot, Lin Wu, Gal Yarin, and Srivastava Akash. Fast and scalable Bayesian deep learning by weight-perturbation in Adam. In Dy Jennifer and Krause Andreas, editors, Proceedings of the 35th International Conference on Machine Learning, volume 80 of Proceedings of Machine Learning Research, pages 2611–2620. PMLR, 10–15 Jul 2018. [Google Scholar]
- [33].Martens James. New insights and perspectives on the natural gradient method. Journal of Machine Learning Research, 21(146):1–76, 2020.34305477 [Google Scholar]
- [34].Carroll SM. Spacetime and Geometry. Cambridge University Press, 2019. [Google Scholar]
- [35].Griffiths DJ. Introduction to Electrodynamics. Cambridge University Press, 2023. [Google Scholar]
- [36].Lin Wu, Dangel Felix, Eschenhagen Runa, Bae Juhan, Turner Richard E., and Makhzani Alireza. Can we remove the square-root in adaptive gradient methods? A second-order perspective. In Salakhutdinov Ruslan, Kolter Zico, Heller Katherine, Weller Adrian, Oliver Nuria, Scarlett Jonathan, and Berkenkamp Felix, editors, Proceedings of the 41st International Conference on Machine Learning, volume 235 of Proceedings of Machine Learning Research, pages 29949–29973. PMLR, 21–27 Jul 2024. [Google Scholar]
- [37].Kirkpatrick James, Pascanu Razvan, Rabinowitz Neil, Veness Joel, Desjardins Guillaume, Rusu Andrei A., Milan Kieran, Quan John, Ramalho Tiago, Grabska-Barwinska Agnieszka, Hassabis Demis, Clopath Claudia, Kumaran Dharshan, and Hadsell Raia. Overcoming catastrophic forgetting in neural networks. Proceedings of the National Academy of Sciences, 114(13):3521–3526, 2017. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [38].Kemker Ronald, McClure Marc, Abitino Angelina, Hayes Tyler, and Kanan Christopher. Measuring catastrophic forgetting in neural networks. Proceedings of the AAAI Conference on Artificial Intelligence, 32(1), Apr. 2018. [Google Scholar]
- [39].Elsayed Mohamed and Mahmood A. Rupam. Addressing loss of plasticity and catastrophic forgetting in continual learning. In The Twelfth International Conference on Learning Representations, 2024. [Google Scholar]
- [40].Kamb Mason and Ganguli Surya. An analytic theory of creativity in convolutional diffusion models. arXiv e-prints, page arXiv:2412.20292, December 2024. [Google Scholar]
- [41].Wang Binxu and Vastola John J.. The unreasonable effectiveness of Gaussian score approximation for diffusion models and its applications. Transactions on Machine Learning Research, 2024. [Google Scholar]
- [42].Vastola John J.. Generalization through variance: how noise shapes inductive biases in diffusion models. In The Thirteenth International Conference on Learning Representations, 2025. [Google Scholar]
- [43].Noether E. Invariante variationsprobleme. Nachrichten von der Gesellschaft der Wissenschaften zu Göttingen, Mathematisch-Physikalische Klasse, 1918:235–257, 1918. [Google Scholar]
- [44].Vastola John J.. Dynamical symmetries in the fluctuation-driven regime: an application of Noether’s theorem to noisy dynamical systems. In NeurIPS 2024 Workshop on Symmetry and Geometry in Neural Representations, 2025. [Google Scholar]
- [45].Tanaka Hidenori and Kunin Daniel. Noether’s learning dynamics: Role of symmetry breaking in neural networks. In Ranzato M, Beygelzimer A, Dauphin Y, Liang PS, and Wortman Vaughan J, editors, Advances in Neural Information Processing Systems, volume 34, pages 25646–25660. Curran Associates, Inc., 2021. [Google Scholar]
- [46].Vastola John J., Gershman Samuel J., and Rajan Kanaka. A variational manifold embedding framework for nonlinear dimensionality reduction. In NeurIPS 2025 Workshop on Symmetry and Geometry in Neural Representations, 2025. [Google Scholar]
- [47].Zenke Friedemann and Gerstner Wulfram. The temporal paradox of hebbian learning and homeostatic plasticity. Current Opinion in Neurobiology, 43:166–176, 2017. [DOI] [PubMed] [Google Scholar]
- [48].Scellier Benjamin and Bengio Yoshua. Equilibrium propagation: Bridging the gap between energy-based models and backpropagation. Frontiers in Computational Neuroscience, Volume 11 - 2017, 2017. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [49].Simmons-Edler Riley, Badman Ryan P., Berg Felix Baastad, Chua Raymond, Vastola John J., Lunger Joshua, Qian William, and Rajan Kanaka. Deep RL Needs Deep Behavior Analysis: Exploring Implicit Planning by Model-Free Agents in Open-Ended Environments. arXiv e-prints, page arXiv:2506.06981, June 2025. [PMC free article] [PubMed] [Google Scholar]
- [50].Khan Mohammad Emtiyaz and Rue Havard. The Bayesian learning rule. Journal of Machine Learning Research, 24(281):1–46, 2023. [Google Scholar]
- [51].Shoji Lucas, Suzuki Kenta, and Kozachkov Leo. Is All Learning (Natural) Gradient Descent? arXiv e-prints, page arXiv:2409.16422, September 2024. [Google Scholar]
- [52].Wibisono Andre, Wilson Ashia C., and Jordan Michael I.. A variational perspective on accelerated methods in optimization. Proceedings of the National Academy of Sciences, 113(47):E7351–E7358, 2016. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [53].Orvieto Antonio and Gower Robert. In Search of Adam’s Secret Sauce. arXiv e-prints, page arXiv:2505.21829, May 2025. [Google Scholar]
- [54].Alshammari Shaden Naif, Hamilton Mark, Feldmann Axel, Hershey John R., and Freeman William T.. A unifying framework for representation learning. In The Thirteenth International Conference on Learning Representations, 2025. [Google Scholar]
- [55].Lecun Y, Bottou L, Bengio Y, and Haffner P. Gradient-based learning applied to document recognition. Proceedings of the IEEE, 86(11):2278–2324, 1998. [Google Scholar]
- [56].Krizhevsky Alex. Learning multiple layers of features from tiny images. 2009. Technical Report. [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
