Abstract
Recent technological advances have enabled measurement of the synaptic wiring diagram, or “connectome,” of large neural circuits or entire brains. However, the extent to which such data constrains models of neural dynamics and function is debated. Here, we develop a theory of connectome-constrained neural networks in which a “student” network is trained to reproduce the activity of a ground-truth “teacher,” representing a neural system for which a connectome is available. Unlike standard paradigms with unconstrained connectivity, the two networks have the same synaptic weights but different biophysical parameters, reflecting uncertainty in neuronal and synaptic properties. We find that a connectome often does not substantially constrain the dynamics of recurrent networks, illustrating the difficulty of inferring function from connectivity alone. However, recordings from a small subset of neurons can remove this degeneracy, producing dynamics in the student that agree with the teacher. Our theory demonstrates that the solution spaces of connectome-constrained and unconstrained models are qualitatively different and determines when activity in such networks can be well-predicted. It can also prioritize which neurons to record to most effectively inform such predictions.
INTRODUCTION
Establishing links between the connectivity of large neural networks and their emergent dynamics is a major goal of theoretical neuroscience. Many studies have attempted to develop methods to infer synaptic connectivity from functional correlations derived from recorded neural activity. However, this “inverse problem” has proven to be challenging and often ill-posed1–5, due to the degeneracy of the space of network connectivities that produce similar dynamics. Such inference is particularly difficult when neural dynamics are low-dimensional or otherwise structured1.
The recent availability of comprehensive synaptic connectome datasets has led to approaches that focus instead on the “forward problem” of predicting neural dynamics from synaptic connectivity. The scale of such datasets has increased rapidly, from the 302 neurons of the nematode C. elegans identified decades ago6 to recently acquired volumes containing entire nervous systems of Drosophila larvae7 and adults8–10, and larval zebrafish11. Several studies have compared connectomes with functional connectivity based on activity correlations between neurons in the resting state or in response to optogenetic perturbations12. This has highlighted striking differences for certain systems13. A complementary line of research has used connectome information to initialize or build explicit priors on the distribution of the parameters of neural network models14;15. In some cases, these models are then optimized to perform computations, and it has been found empirically that such biological constraints sometimes yield models with improved abilities to predict neural data16–19. However, the ill-posedness of the inverse problem and lack of one-to-one correspondence between structure and function call into question the reliability of such predictions.
A major challenge for connectome-constrained models is uncertainty in biophysical parameters that affect neural dynamics. Connectomes generated from electron-microscopy imaging provide information on structural connections, neurotransmitter identities of chemical synapses10;20, and connection strengths estimated by synapse count21 or volume22. However, other biological processes are undetermined, such as the neuromodulatory environment, existence of electrical synapses, and functional properties of individual neurons and synapses23. Changes in such parameters have previously been shown to produce dramatic alterations in network activity24–27.
In this study, we develop a theory of the solution spaces of networks with specified synaptic weights but unmeasured and heterogeneous single-neuron biophysical parameters28;29. We use a “teacher-student” paradigm in which the activity of a “student” network is trained to reproduce the activity of a “teacher” network. The teacher is a synthetic model that represents ground truth, analogous to biological circuits for which a connectome is available and from which we can record activity. Unlike previous theories in which the student and teacher neurons have the same input-output function and synaptic weights are trained1;30, here the two networks have the same weights but their biophysical properties differ a priori.
We find that training a connectome-constrained student network to generate the task-related readout of the teacher does not always produce consistent dynamics in the teacher and student. Multiple combinations of single-neuron parameters, each producing different activity patterns, can equivalently solve the same task. However, when connectivity constraints are combined with recordings of the activity of a subset of neurons, this degeneracy is broken. The minimum number of recordings depends on the dimensionality of the network dynamics, not the total number of neurons. This contrasts with student networks whose connectivity is unconstrained, which always display degenerate solutions. Interestingly, even when neural activity is well-reconstructed, single-neuron parameters are often not recovered accurately, suggesting that some combinations of parameters are “stiff,” with strong effects on neural dynamics, while others are “sloppy,” with weak effects. Our qualitative predictions hold across a variety of simulated networks and networks constrained by true connectomes from invertebrates and vertebrates. Our theory can also rank neurons that should be recorded with higher priority to maximally reduce uncertainty in network activity, suggesting approaches that iteratively refine network models using neural recordings.
RESULTS
Teacher-student recurrent networks
To explore how a connectome constrains the solutions of neural network models, we studied a teacher-student paradigm31;32: a recurrent neural network (RNN) that we call the teacher is constructed, and the parameters of a student RNN are adjusted to mimic this teacher. The teacher is used as a proxy for a neural system whose connectome has been mapped and whose output or neural activity can be recorded. To develop our theory, we will begin by examining synthetic teacher networks whose activity and function we specify. Later, we will consider teacher networks derived from empirical connectome data.
Both teacher and student are composed of firing rate neurons, in which the activity of neuron is described by a continuous variable (see Methods for details). The activity is a nonlinear function, which we call the activation function, of the input current received by the neuron, and depends on a set of single-neuron parameters. For instance, if we describe this function using parameters and for neuron ’s gain and bias, its activity is given by
| (1) |
where is a nonlinear function. The network dynamics follow:
| (2) |
where is the synaptic weight from neuron to neuron , and is the time-varying external input received by neuron . For connectome-constrained networks, we begin by assuming that both the presence or absence of a connection between neurons as well as the strengths of these connections are known, and thus is the same for both teacher and student. Additionally, we assume that the external inputs and initial state are the same for teacher and student (see Discussion).
Note that the number of unconstrained parameters in the student network scales differently depending on whether single-neuron parameters or connectivity parameters are fixed. There are free synaptic weight parameters if the connectivity is unspecified, as in prior studies of teacher-student paradigms31;32. On the other hand, for connectome-constrained networks, the number of unconstrained parameters is proportional to . For example, when we parameterize neurons’ activation functions with gains and biases, as in Equation (1), there are unknowns.
Student network constrained by task output
We first asked whether teacher and student networks that share the same synaptic weight matrix exhibit consistent solutions when the student is trained to reproduce a task performed by the teacher (Fig. 1). Because we are interested in whether connectivity constraints yield mechanistic models of the teacher, we measure the consistency of solutions using the similarity of the activity of neurons in the teacher and those same neurons in the student. Such a direct comparison is possible because the connectome uniquely identifies each individual neuron. We also measure the similarity of teacher and student single-neuron parameters. We refer to the dissimilarity between teacher and student activities or parameters as the “error” associated with each respective quantity. We note that our notion of similarity between teacher and student is more precise than requiring similarity of collective dynamics as measured through dimensionality reduction methods, such as principal components analysis. Indeed, matching such dynamics can be accomplished by recording a small number of neurons without access to a connectome33;34.
Figure 1: Task-trained networks with the same connectivity.

A A teacher RNN is trained to generate two different readout sequences in response to input pulses that produce two different patterns of activation (gray and black).
B Properties of the teacher RNN. The teacher RNN has heterogeneous single-neuron parameters (gains and biases of activation functions, left), sparse connectivity with connection probability (right), and neurons connect through either excitatory (red) or inhibitory (blue) synapses.
C Student networks with the same synaptic weights as the teacher are trained to produce the teacher’s output. Error in the readout (training loss, MSE) as a function of training epoch. Each colored line corresponds to a different student network.
D Error (mismatch in neural activity) between teacher and student RNNs. Gray line, for reference, corresponds to the average error in activity when the student reproduces the teacher’s activity but with shuffled neuron identities.
E Error in gains and biases vs. training epoch.
F Readout of teacher and student networks after training, for the two trial types (top and bottom). Teacher and student networks both solve the task.
G Neural activity of an example excitatory (left) and inhibitory (right) neuron. Teacher and student neurons exhibit different single-neuron dynamics.
We built a teacher network that performs a flexible sensorimotor task. Specifically, the network implements a variant of the cycling task35, which requires the production of oscillatory responses of different durations in response to transient sensory cues (Fig. 1 A; see Methods). In the network, firing rates are a non-negative smooth function of the input currents, and the unknown single-neuron parameters are the gains and biases (Fig. 1 B left). The synaptic weight matrix is sparse and neurons are either excitatory or inhibitory (Fig. 1 B right).
We trained multiple students to generate the same readout as the teacher. Each student is initialized with different gains and biases before being trained via gradient descent. Trained networks successfully reproduce the teacher’s readout (Fig. 1 C, F). However, the error in the neural activity of the student, compared to the teacher, increases over training epochs (Fig. 1 D). As a baseline, we computed the error of a student whose neurons match the activities of all neurons in the teacher, but with shuffled identities (gray line, Fig. 1 D). In this baseline, the manifold of neural activity is the same in teacher and student, but not the activity of single neurons. In all networks, the error in activity after training remains above this baseline, indicating that training does not produce a correspondence between the function of individual teacher and student neurons. Examining the activities of individual neurons shows that neuronal dynamics across different student networks are highly variable, and all students differ from the teacher (Fig. 1 G).
Finally, we examined the error in single-neuron parameters between teacher and student (Fig. 1 E). The error in gains varies little over training and is comparable to a randomly shuffled baseline. The error in biases grows slightly but remains within the same order of magnitude as the baseline.
We conclude that knowledge of synaptic weights and task output is not always enough to predict the activity of single neurons in recurrent networks. For the task we considered, there is a degenerate space of solutions, with different combinations of single-neuron gains and biases, that solve the same task. There may be scenarios for which this degeneracy is reduced, such as small networks optimized for highly specific functions, or networks trained on complex or high-dimensional task spaces (see Discussion). Nonetheless, our results show that, even with connectivity constraints, task-optimized neural dynamics are, in general, highly heterogeneous.
Student network constrained by activity recordings
We next asked whether these conclusions change if, instead of recording only task-related readout activity, we record the activity of a subset of neurons in the teacher network. We use to denote the number of recorded neurons. Students are trained to reproduce this recorded activity, which provides additional constraints on the solution space (Fig. 2 A,B). The recording of subsampled activity in the teacher is analogous to neural recordings in imaging or electrophysiology studies, where only a subset of neurons are registered. We trained two types of student networks: students that have access to the teacher connectome, and students that are not constrained in connectivity. For connectome-constrained students (Fig. 2 C,E), single-neuron parameters of both recorded and unrecorded neurons are unknown and therefore trained. For students with unconstrained connectivity, synaptic weights are trained instead. In this case, the single-neuron parameters of the student are set equal to those of the teacher so that the networks differ only in synaptic weights (Fig. 2 D,F). Additionally, since there is no direct map between unrecorded neurons in the teacher and the student when the connectome is not known, after training we search for the mapping between student and teacher neurons that minimizes the mismatch in unrecorded activity at each training epoch (see Methods).
Figure 2: Predicting activity of unrecorded neurons when the activity of a subset of the network is observed.

A The student RNN is trained to mimic the activity of recorded neurons in a teacher RNN.
B Error in recorded activity (loss) vs. training epoch for students with trained single-neuron parameters (left) and students with trained synaptic weights (right). Lines correspond to different numbers of recorded neurons and show mean over 10 random seeds. Error bands in all panels indicate ±SEM. All students successfully reproduce the recorded activity of the teacher after training.
C Left: Error in activity of the unrecorded neurons vs. training epoch. Right: Error in unrecorded neuronal activity after training, as a function of number of recorded neurons . Smaller dots correspond to each of the 10 trained students. Error is substantially reduced when recording from neurons. The error corresponding to zero recorded neurons (black dots) is the error of the student network prior to training, with random single-neuron parameters. Gray line denotes shuffled baseline as in Fig. 1 D.
D Analogous to Panel C, but training synaptic weights instead. The error in the activity of unrecorded neurons remains high across values of . The -dependence is a consquence of the procedure of matching neurons across teacher and student (see Methods).
E Error in gains and biases vs. training epochs. Left: Parameters of recorded neurons. Right: Parameters of unrecorded neurons.
F Analogous to panel E for synaptic weights of recorded neurons (left) and unrecorded neurons (right).
We found that both connectome-constrained and unconstrained students are able to mimic the activity of the recorded teacher neurons with small errors (Fig. 2 B, the teacher has neurons). We then asked whether this holds for the unrecorded neurons. When the connectivity is provided (Fig. 2 C), the error for unrecorded neurons is reduced to values comparable to the error for recorded neurons when more than neurons are recorded (example task outputs are shown in Extended Data Fig. 1). In comparison, when training the synaptic weights (Fig. 2 D), unrecorded neurons’ activities are not recovered substantially better than baseline even when a majority of neurons is recorded. Thus, connectome-constrained, but not unconstrained, networks produce consistent solutions when is large enough.
We then assessed whether the students’ parameters converge to those of the teacher. For connectome-unconstrained students, the error in synaptic weights remains high, both for connections between recorded and unrecorded neurons (Fig. 2 F). We may expect this to occur given that the activity of unknown neurons in these networks is not well-predicted (Fig. 2 D). More surprisingly, errors in the single-neuron parameters of connectome-constrained networks also remain high, even when the activity of unrecorded neurons is well-predicted (Fig. 2 E). We did not find qualitative differences in the behavior of single-neuron parameters for recorded and unrecorded neurons.
So far, we focused on a teacher whose neural activity is primarily generated through recurrent interactions, triggered by brief external pulses. We further explored whether similar results hold in networks driven by a time-varying external input (Extended Data Fig. 1). Additionally, we systematically varied the distributions of gains and connection sparsity (Extended Data Fig. 1). In all these networks, the qualitative dependence of the error on was unchanged. Nevertheless, the error in unrecorded neural activity prior to training is different in networks with strong inputs or weak recurrent connections. Unlike in Fig. 2, where the error prior to training is comparable to a baseline with randomly shuffled neuron identities, the error for strongly input-driven networks lies below this baseline even before training. Thus, while certain features of neural activity may be predictable even with random parameters when the input is known, improving upon this initial baseline through training requires sufficiently many recorded neurons.
In summary, connectome-constrained networks are able to predict the activity of unrecorded neurons when further constrained by the activity of enough recorded neurons. In contrast, networks without a connectome constraint do not predict unrecorded activity. Nevertheless, in all cases, the unknown parameters are never precisely recovered, suggesting that multiple sets of biophysical parameters lead to the same neural activity.
Required number of recorded neurons is independent of network size
What features of a connectome-constrained RNN determine how many recorded neurons are required to predict unrecorded activity? We considered two alternatives: the required number is a fixed fraction of the total number of neurons in the network, or the number is determined by properties of the network dynamics. The former alternative would pose a challenge for large connectome datasets.
To disambiguate these two possibilities, we examined a class of teacher networks whose population dynamics are largely independent of their size . We generated networks with specific rank-two connectivity that autonomously generate a stable limit cycle36 (Fig. 3 A and Methods). In these networks, the currents received by each neuron oscillate within a two-dimensional linear subspace, independent of (Fig. 3 B).
Figure 3: Prediction of unrecorded neurons’ activity depends on dimensionality, not network size.

A Set of teacher RNNs with variable network size N but fixed rank of the synaptic weight matrix (see Methods).
B Teachers of different size produce the same low-dimensional dynamics. Left: Dynamics projected on the top two principal components (PCs). All RNNs generate a limit cycle largely constrained to a 2D linear subspace. Right: Variance screeplot.
C Error in the activity of unrecorded neurons, after training. We measured the correlation distance between activity in the teacher and student. Plot shows empirical average and SEM for each network size (10 networks per condition).
D Set of teacher RNNs with variable network size and random full-rank synaptic weights.
E Left: Networks generate high-dimensional chaotic dynamics. Sample activity of four units for networks of different sizes and a time window of . Right: Variance scree plot of the recordings. Larger networks generate higher dimensional dynamics.
F Error in the activity of unrecorded neurons after training. Larger networks require recording from a larger number of neurons to predict unrecorded activity. Data are presented as mean ± SEM over 10 networks, as in panel C. Inset: Number of neurons needed to predict unrecorded activity above a certain threshold (set to 0.2; dotted line), as a function of network size.
We found no difference in a plot of error in unrecorded activity against number of recorded neurons , for networks of different sizes (Fig. 3 C), suggesting that accurate predictions can be made when recording from few neurons, even in large networks. Examining more closely the dependence of the error on , we observed that when , the student produces oscillatory activity with the same frequency as the teacher, but the activity of unrecorded neurons exhibits consistent errors at particular phases of the oscillation (Extended Data Fig. 2). In contrast, when , errors in recorded and unrecorded neurons are comparably small.
This led us to hypothesize that the number of recorded neurons required to accurately predict neural activity scales with the dimensionality of the neural dynamics, not the network size. This would explain why networks with widely varying sizes but similar two-dimensional dynamics exhibit comparable performance (Fig. 3 C). To further test this hypothesis, we studied a setting in which we trained students to mimic another class of teacher networks, strongly-coupled random networks37 (Fig. 3 D). In such networks, activity is chaotic, and unlike low-rank networks (Fig. 3 B), the linear dimensionality of the dynamics grows in proportion to 38, a dependence we verified for the time windows we considered (Fig. 3 E). In this case, the required number of recorded neurons also grows proportionally with (Fig. 3 F). Together, these results suggest that recording from a subset of neurons, on the order of the dimensionality of network activity, is sufficient to predict unrecorded neural activity. Later, we will show that this numerical result is consistent with the predictions of an analytical theory.
Robustness to model mismatch
Thus far, we have considered teacher and student networks that belong to the same model class of firing-rate networks with parameterized activation functions and connectivity. However models based on experimental data will possess some degree of “model mismatch” due to unaccounted or incorrectly parameterized biophysical processes. Moreover, errors in synaptic reconstruction and inter-individual variability in connectomes imply that synaptic weight estimates may also be imprecise39. In this section, we examine whether our qualitative results hold when teacher and student exhibit model mismatch.
We used the same teacher as in Figs. 1 and 2. To study the case of mismatch in activation function (Fig. 4 A), we parameterized the activation function with , which controls the smoothness of the rectification, and used different values of for student and teacher (Fig. 4 B). Larger mismatch increases the error in both recorded and unrecorded activity (Fig. 4 C). The effect is strongest in an extreme case of very small student , for which very little rectification occurs. This makes it difficult for the student to match even the recorded activity of the teacher (Fig. 4 C, Extended Data Fig. 3). Nevertheless, up to a considerable mismatch, there is a steep decrease in the error in unrecorded activity as more neurons are recorded.
Figure 4: Model mismatch between teacher and student.

A Mismatch in activation functions of teacher and student neurons.
B The activation function is a smooth rectification but with different degrees of smoothness, parameterized by . Teacher RNN from Fig. 2.
C Errors in the activity of recorded (left) and unrecorded (right) neurons for different values of model mismatch between teacher and student. Across a large range of mismatch we observe a decrease of the error in unrecorded activity when .
D Mismatch in synaptic weights between teacher and student, mimicking errors in connectome reconstruction.
E Eigenvalues of the teacher and student synaptic weight matrices, for different levels of mismatch.
F Errors in the activity of recorded (left) and unrecorded (right) neurons for different levels of mismatch in the synaptic weights.
Can model mismatch arising from single-neuron properties be compensated by allowing the synaptic weights to be trained, which introduces additional free parameters? We examined a student with activation function mismatch, and a synaptic weight matrix that was initialized equal to that of the teacher, but then trained (Extended Data Fig. 3). This performed worse than training single-neuron parameters, arguing against the feasibility of this approach. An alternative approach is to increase the number of single-neuron parameters. For instance, when is trained together with gains and biases, the error in unrecorded activity is similar to the case without mismatch (Extended Data Fig. 3). We conclude that parameterizing uncertainty in activation function is important for dealing with this form of model mismatch.
We next considered mismatch between teacher and student connectomes. To simulate such errors, we added Gaussian noise to the strengths of existing connections and added spurious connections with probability (Fig. 4 D; see Methods). The resulting corrupted synaptic weight matrix was used by the student. Noise in the synaptic weight matrix shifts its eigenvalues (Fig. 4 E) and modifies the corresponding eigenvectors. Trained students exhibit smooth increases of the error in recorded and unrecorded activity as this noise is increased (Fig. 4 F). However, we again found a steep decrease of the error in unrecorded neural activity with , suggesting that this qualitative behavior is not overly sensitive to connectome reconstruction errors.
Teacher networks constrained by empirical connectomes
So far, we have examined synthetic teachers, whose connectivity statistics and functional properties may differ from those of biological networks. We next study teachers whose synaptic weights are directly determined by empirical connectome datasets. We modeled three neural circuits for which a ground-truth connectome is available and whose function has been characterized: the premotor-motor system in the ventral nerve cord of larval Drosophila16, the heading direction system in the central complex of adult Drosophila9;40, and the oculomotor neural integrator in the hindbrain of larval zebrafish41.
When larval Drosophilae are engaged in forward or backward locomotion, recurrently connected premotor neurons in the ventral nerve cord drive motor neurons to produce appropriately timed muscle activity (Fig. 5 A). Motor neurons in each body segment are segregated into functional groups whose sequences of activation differ across the two behaviors (Fig. 5 F). A previous study showed that a connectome-constrained RNN recapitulates features of motor and premotor neuron activity when trained to produce such sequences in the A1 and A2 body segments16. We used such a model as a connectome-constrained teacher, whose 178 premotor neurons produce appropriately timed activity in 52 motor neurons (see Methods). Student networks comprising the premotor circuitry were then trained to approximate recorded teacher activity. We found that the error in unrecorded activity is reduced when approximately 10 neurons are recorded (Fig. 5 B). When few neurons are recorded, the error in activity is comparable to a network with randomly chosen single-neuron parameters (two recorded neurons, Fig. 5 C left). Recording from more neurons dramatically improves the prediction (Fig. 5 C right), which is qualitatively similar to the results of the synthetic teacher network (Fig. 2 C).
Figure 5: Prediction of neural activity in networks with empirical connectome constraints.

A Left: Diagram of premotor (PMNs) and motor neurons (MNs) in segments A1 and A2 of the Drosophila larva. The model synaptic weights are determined by a connectome, as in Zarin et al. (2019)16. Right: Activity of MN readouts during forward (FWD, top) and backward (BWD, bottom) crawling.
B Error (based on the average single-neuron correlation between teacher and student) in unrecorded activity of PMNs as a function of the number of recorded neurons. Black dot indicates error before training, i.e. the error before recording any neurons. Arrows indicate number of recorded neurons for which example traces are shown in the adjacent panel.
C Example activity traces of two unrecorded PMNs in the teacher network, and in the student network before and after training, when two neurons are recorded (left) and when 24 neurons are recorded (right). Traces are normalized by their maximum value.
D Left: Synaptic weight matrix for neurons of the central complex (CX) of adult Drosophila, based on the hemibrain connectome9. Right: Input and target output of EPG neurons, which are arranged along the ellipsoid body according to their tuning to heading direction angle (top). On each trial, the input is a transient bump of activity presented at a random orientation. The target output requires this orientation be held in a persistent activity state.
E Similar to B for the CX model.
F Similar to D, example traces of one unrecorded neuron when two different presented stimuli (one preferred, one non-preferred). Traces are normalized by the mean firing rate of each neuron across stimuli.
G Left: Diagram of synaptic wiring in the zebrafish brainstem, based on Vishwanathan et al. (2024)41. Right: Traces of activity of three example neurons. There is a slow mode of activity which integrates eye velocity and is required for the oculomotor reflex.
H Similar to B, E for the oculomotor integrator model.
I Similar to C, F for the neurons shown in G, right.
Next, we studied the heading direction system in the central complex of adult Drosophila. This system has been the subject of numerous recent theoretical analyses, most of which examined models with idealized connectivity rather than directly incorporating connectome data40;42;43. We modeled a circuit reconstructed in the hemibrain dataset9 comprising 153 neurons grouped into four cell types: the putatively excitatory E-PG, P-EN, and P-EG neurons, and the putatively inhibitory Delta7 neurons (Fig. 5 D). The 46 E-PG neurons encode heading orientation and are arranged along a ring in the ellipsoid body based on their angular tuning. Recurrent connections among EP-Gs and other cell types form a stable “bump” of neural activity representing heading angle, consistent with “ring attractor” dynamical models44. We therefore constructed a teacher network in which EP-G neurons maintained a bump representing a heading encoded by a brief stimulus (Fig. 5 D; see Methods). Student networks without access to recordings generated neural activity different from the teacher (Fig. 5 E, black dots). In particular, these students did not behave as ring attractors, demonstrating that the central complex connectivity alone does not guarantee stable attractor dynamics (Fig. 5 F). However, recording from a handful of neurons was enough to place the system in the correct dynamical regime and accurately predict the activity of unrecorded neurons (Fig. 5 F).
Finally, we studied the oculomotor integrator in the hindbrain of larval zebrafish. This system persistently tracks eye position by integrating eye motor commands. The integration is supported by strong recurrent connections that produce a “line attractor” in neural activity space. Such dynamics were previously modeled with a connectome-constrained linear RNN41 (Fig. 5 G, see Methods). We used this network as the teacher and then trained the gains of student networks with the same synaptic weights. Although a random initialization of gain parameters did not produce the slow timescale necessary for accurate integration, recording from a few neurons substantially reduced the error in activity (Fig. 5 H). This is consistent with the results of Vishwanathan et al. (2024)41, who adjusted a global gain parameter to produce a slow timescale.
The weight matrices of empirical connectomes and synthetic teacher-student networks (Figs. 2, 3) may exhibit statistical differences due to the level of sparsity, heterogeneity in the number and strength of synaptic connections, and other higher-order structure. However, in each of these examples, the qualitative phenomena present in synthetic teacher-student networks are recapitulated. Recording from a number of neurons determined by the dimensionality of the teacher activity (see Extended Data Fig. 4)— a handful for the one-dimensional line attractor or two-dimensional ring attractor dynamics and ~10 for more complex sequential activity — produces consistent dynamics between teacher and student.
Linear network model
We developed an analytic theory of our connectome-constrained teacher-student paradigm. The theory aims to explain, first, how the teacher and student produce the same activity despite different single-neuron parameters, and second, the conditions under which the student’s activity converges to that of the teacher.
We begin with a simplified linear model and later relax our assumptions: the teacher and student RNNs have linear single-neuron activation functions, the only unknown single-neuron parameters are the biases , and the synaptic weight matrix has rank (Fig. 6 A). This rank constraint implies that recurrent neural activity is confined to a -dimensional subspace of the -dimensional neural activity space. We focus on the network’s steady-state activity at equilibrium, which depends linearly on the biases:
| (3) |
where we have defined .
Figure 6: Linear teacher-student model.

A Left: The activity of a neuron is a linear mapping , which depends on the synaptic weight matrix , of the single-neuron parameters . Right: Neuronal activation functions are linear with heterogeneous biases.
B Singular values of the synaptic weight matrix, which is random and has rank (gray line).
C Errors in activity and biases as a function of the number of recorded neurons.
D Single-neuron biases evolve over training through gradient descent. Parameter modes are described as stiff or sloppy based on the effect of changes along each mode near the optimal solution.
E Singular value decomposition of the mapping determines stiff and sloppy parameter modes. Stiffer modes are learned more quickly. Gray line at corresponds to the rank of the synaptic weight matrix.
F Effective singular value decomposition when recording from a subset of neurons. Inset shows the maximum angle between the stiffest modes and the sub-sampled parameter modes. Different colors correspond to different numbers of recorded neurons (as shown in the inset).
G-H Evolution of errors in activity and biases for and , for 10 different initializations of parameters. Error in biases is projected along one stiff (1st) and one sloppy (50th) parameter mode.
Although we focus here on equilibrium activity, time-dependent trajectories also yield a linear relation between activity and single-neuron parameters (see Methods for the time-dependent derivation). For the same reason, we also assume no external input to each neuron (). This linear relation between single-neuron parameters and activity, which underpins the mathematical tractability of the simplified model, is a consequence of the linear network dynamics and the additive influence of the bias parameters. Choosing multiplicative gains as the unknown single-neuron parameters, for instance, would produce a nonlinear relation.
The student is trained using gradient descent updates to the single-neuron parameters. In the limit of small learning rate , the learning trajectory in parameter space can be expressed in continuous time (with proportional to training epoch) as:
| (4) |
where is the number of recorded neurons. Using these learning dynamics, we can analytically calculate the expected error in recorded and unrecorded activity and in single-neuron parameters (Fig. 6 C; see Methods). This reveals a transition to zero error in the activity of unrecorded neurons when , the rank of the synaptic weight matrix (Fig. 6 C, gray line). There are, however, large errors in single-neuron parameters (Fig. 6 C, red line) even when the activity of the full network is accurately recovered.
To understand these results, we analyzed the properties of the loss function, which describes how the difference in activity between teacher and student depends on single-neuron parameters. We differentiate the loss function for the full network, which is determined by errors in both recorded and unrecorded neural activity, from the loss function for the recorded neurons, which is the function optimized during training. These loss functions are convex, as illustrated in Fig. 6 D. The minima are surrounded by a valley-shaped region of low loss (Fig. 6 D right). We refer to directions for which the loss changes quickly or slowly as “stiff” or “sloppy” parameter modes, respectively45. Stiff modes both have the greatest effect on the loss and are learned most quickly. Each mode’s degree of stiffness is determined by the corresponding singular value of the matrix (see Methods). A mode is infinitely sloppy when its associated singular value is zero, implying that parameter differences between teacher and student along that mode produce no differences in neural dynamics.
The parameter modes that affect the recorded activities and thus the loss function for the recorded neurons are determined by (the submatrix of containing the rows corresponding to these neurons), whose stiff and sloppy modes are generally different from those of the fully sampled matrix (Fig. 6 E vs. F). Recording from a subset of neurons introduces additional modes with zero singular value when , since has at most non-zero singular values. The stiff modes of the loss function of the recorded activity will also, typically, not be fully aligned with those of the fully sampled system (Fig. 6 F, inset), leading to errors in prediction.
To illustrate these results, we plotted the error in single-neuron activity and biases for below and above the critical number (Fig. 6 G, H). When recording from few neurons, the error in unrecorded neurons remains high (Fig. 6 G left). The error in biases quickly converges to a small value along the stiffest mode, while it barely changes for sloppy modes (Fig. 6 G right). The stiffest mode of the subsampled network is not completely aligned with the stiffest mode of the fully sampled network, explaining why it converges to a small but nonzero value. Only when more neurons are recorded does the error in unrecorded activity, and along the stiffest parameter mode, converge to zero (Fig. 6 H left).
The simplified model demonstrates that specific patterns of single-neuron parameters determine the error between teacher and student. Stiff parameter modes are learned, while sloppy modes are not. The number of stiff parameter modes is bounded by the rank of and thus the dimensionality of neural activity, and recording from increasingly many neurons provides increasingly many constraints on this activity. When enough neurons are sampled, the stiff modes for the recorded neurons align with the stiff modes for the full network, leading to correct prediction of unrecorded activity. These conclusions do not rely on the choice of gradient descent as a learning algorithm, as an analysis using linear system identification methods yields the same conclusions (see Supplementary Note). We also note that we have assumed here that the loss function is determined by the difference in recorded neural activity. However, similar conclusions would be reached if it were determined by other linear projections of activity, such as projections onto task-related dimensions (see Discussion).
Loss landscape
We next generalized our theory to nonlinear networks. To facilitate analysis, we studied a class of low-rank RNNs whose activity can be understood analytically36;46;47. We focused on a teacher network with neurons and a nonlinear, bounded activation function. Each neuron is parameterized only by the gain parameter. We designed the network’s synaptic weight matrix to be rank-two, with two different subpopulations. For this network, there are only two stiff parameter modes, the average single-neuron gain for each subpopulation (Fig. 7 A). We set the weights of the teacher to generate two different pairs of nontrivial fixed points, and we recorded activity as the neural dynamics approached one of these fixed points (Fig. 7 B).
Figure 7: Loss landscape in nonlinear networks.

A We study a rank-two RNN with two populations. Neurons in each population share the same gains and connectivity statistics.
B Dynamics of the target network. Left: Phase-space of the two-dimensional latent variables. Right: Activity as a function of time for 20 sampled neurons. For illustrative purposes, neurons 1 and 2 are selected based on their alignment with the two latent variables.
C The loss function of the full network depends only on the gains of each population, and . White dot indicates the parameters of the teacher RNN.
D Loss function when recording the activity of neuron 1 (left) or neuron 2 (right). Blue and red squares correspond respectively to solutions where the training loss is close to zero.
E Target trajectory (black) and dynamics of the teacher RNN. Blue and red trajectories correspond to the solutions found in D.
F Predicted activity for neurons 1 and 2 for the solutions found in D. Left: Error in the activity of the recorded neuron (neuron 1) is small, while error for the unrecorded neuron (neuron 2) is large. Right: Similar to left, but when neuron 2 is recorded and neuron 1 is unrecorded.
G Trained nonlinear RNN from Fig. 1.
H Average squared error in parameters projected on the different stiff and sloppy parameter modes. The stiff and sloppy dimensions are determined by approximating the full-sampled loss function around the teacher’s values (see Methods). Average over 10 realizations.
Since the parameter space is two-dimensional, we can visualize the loss function for the full network across a grid of parameters (Fig. 7 C). The function has a single minimum, similar to the linear model. However, due to the nonlinearity, the function is non-convex (contour lines are not convex in Fig. 7 C), and the curvature for parameter values away from the global minimum is different than at the minimum. Despite this non-convexity, gradient descent on this fully sampled loss function will still approach the single minimum.
We next visualized the loss function for the activity of one recorded neuron (Fig. 7 D). We repeated this for two different choices of the single recorded neuron, each of which exhibited distinct dynamics (black lines, Fig. 7 B). For these loss functions, there is an additional sloppy mode that is not present in the fully sampled loss (black valleys in Fig. 7 D). These results are similar to those of the linear case, although due to the nonlinearity, the sloppy modes correspond to curved regions in parameter space.
The sloppy mode is different for each of the two recorded neurons. When running gradient descent on these subsampled loss functions, randomly initialized parameter values will evolve toward the dark regions of Fig. 7 D, for example toward the blue dot when neuron 1 is sampled or toward the red dot when neuron 2 is sampled. However, both of these two solutions produce high error in unrecorded activity (Fig. 7 E,F). This mismatch in unrecorded activity occurs because recording from a single neuron constrains activity along only one dimension of the two-dimensional activity space defined by the rank-2 synaptic weight matrix (Fig. 6 E).
To test whether the same insights also apply to nonlinear networks with high-dimensional parameter spaces, we computed the stiff and sloppy modes of the fully sampled loss function in the network of Figs. 1 and 2. We approximated the loss function in parameter space to second order at the optimum. We then projected the average error in parameter space, before and after training, along the estimated stiff and sloppy modes (Fig. 7 G–H). When few neurons are recorded, the average changes in parameter space before and after training are not aligned with the stiff modes of the loss function. However, when recording from many neurons, there is a large decrease in error along the estimated stiff modes while the error along sloppy modes barely changes, as predicted by our theory. Thus, a second-order approximation of the non-convex loss function qualitatively describes the behavior of gradient descent. Some other effects of non-convexity, however, cannot be explained by the linear theory, for instance the growth in errors in the bias parameters over the course of training (Figs. 1 E and 2 E).
We conclude that the qualitative behavior of the linear model holds for the nonlinear networks studied in previous sections. Specifically, when the loss function is determined by recordings of a small number of neurons, the parameter modes become sloppier on average, and new sloppy parameter modes are added that do not align with those of the fully sampled loss function.
Optimal selection of single neurons
So far, recorded neurons have been selected randomly from the teacher network. As we have seen, different sets of recorded neurons define different loss functions and gradient descent dynamics, suggesting the possibility of selecting recorded neurons to minimize the expected error in unrecorded activity (Fig. 8 A). Specifically, we aim to select recorded neurons to maximize the alignment of stiff modes of the subsampled loss and those of the fully sampled loss function.
Figure 8: Optimal selection of recorded neurons.

A Recording from specific subsets of neurons (right) in the teacher RNN leads to different performance.
B We linearized the mapping from changes in single-neuron parameters to changes in neural activity.
C Teacher RNN with linear single-neuron activation functions, unknown biases, and synaptic weight matrix with rank (as in Fig. 6).
D Error in activity of unrecorded neurons as a function of number of recorded neurons . Lines correspond to theoretical prediction, dots to numerical simulation (mean ± SEM). We selected neurons following the estimated best ranking (red), 5 different random rankings (black), and the worst ranking (blue).
E Error in recorded neurons for the same networks.
F - H Analogous to C-E but for a nonlinear network, data shows mean ± SEM. The teacher is the RNN from Fig. 2. Single-neuron parameters are both gains and biases. The linearization in the mapping from parameters to activity assumes homogeneous single-neuron parameters (see Methods). Note the linear scale for the error in panel G, which highlights the errors when few neurons are recorded.
In the simplified linear model, subsampling neurons corresponds to selecting rows of the matrix that relates single-neuron parameters to activity (Fig. 8 B). In this case, it is possible to exactly determine which neurons are most informative to record. The most informative neuron is the one whose corresponding row overlaps most with the weighted left singular vectors of (see Methods). The second most informative neuron is the one whose row overlaps most with the weighted left singular vectors of projected onto the space orthogonal to the previously selected neuron’s row, and so on. It is also possible to define the least informative sequence of recorded neurons by minimizing, rather than maximizing, these overlaps. We compared the error in unrecorded activity for the most and least informative sequence of selected neurons as well as random selection, finding that the optimal strategy indeed improves the efficiency of training (Fig. 8 D).
For nonlinear networks, the mapping between parameters and network activity is also nonlinear and depends on the unknown parameters of the teacher (see Methods). As a result, the globally optimal sequence cannot be determined a priori. Nevertheless, the mapping between parameters and activity can be linearized based on an initial guess of the single-neuron parameters and then iteratively refined. In practice, we found that linearization works well for nonlinear networks, with the optimal selection strategy dramatically reducing the error compared to random selection, especially when there are few recorded neurons. For the network studied in Fig. 2, the error using the best 10 predicted neurons is 60% smaller than random selection (Fig. 8 G).
The singular vectors used to determine which neurons are most informative depend on the global connectivity structure and cannot be exactly reduced to any single-neuron property. Such properties, including in-degree, out-degree, average synaptic strength, or neuron firing rate, may be correlated with the singular value decomposition score developed here, but are not guaranteed to be good proxies for informativeness. This argues for the use of models like those studied here to guide the selection of recorded neurons.
DISCUSSION
Building connectivity-constrained neural network models has become increasingly viable as the scale of connectome datasets has grown. Our theory cautions against over-interpreting such models when they are insufficiently constrained (Fig. 1), but also shows that correctly parameterized models paired with sufficiently many neural recordings can provide consistent predictions (Figs. 2, 5). This consistency is a consequence of the qualitatively different solution spaces associated with connectome-constrained and unconstrained models (Fig. 2 C,D, Fig. 7). The theory also suggests that models can be used to inform targets for physiological recordings (Fig. 8).
Challenges for connectome-constrained neural networks
We have studied the properties of connectome-constrained neural networks using simulations of neural activity based on synthetic network models and three different connectomics datasets9;10;41. Our results suggest that the “forward problem” of predicting neural activity using a connectome is not as ill-posed as the corresponding “inverse problem” studied previously1. However, although we demonstrated that this result is robust to model mismatch and inaccuracy in synaptic reconstruction (Fig. 4), it is likely that for some neural systems the degree of model mismatch is too severe. Such systems likely include those largely driven by unmodeled processes such as the effects of neuropeptides or gap junction couplings13;23. Moreover, systems for which the firing rate models described here are a poor match, such as systems that operate based on spike synchrony rather than rate codes48, highly compartmentalized interactions49, or dynamics of specific ion channels50, may be out of reach of the present approach. Nonetheless, our results establish that the dynamics in RNNs with order , rather than , unknown parameters, can be accurately predicted.
We also assumed in all our analyses that time-varying external inputs, together with the initial state, are known. Consistent with this, recent work using connectome information to infer function has focused on regions close to the sensory periphery19;51, where input statistics are better characterized. We do not expect connectomes to provide substantial constraints on strongly input-driven neural activity when inputs are not controlled.
Relatedly, for our studies of empirical connectomes (Fig. 5), we modeled systems for which detailed descriptions of the function of individual cell types exist16;41;42. This was necessary to generate realistic teacher activity, as recordings of neural activity aligned to whole-brain connectomes are not yet available for the systems we studied. We did not apply our theory to C. elegans recordings as it has been argued that chemical synapses are not predictive of functional interactions13.
We note that the parameters in our models may reflect state-dependent modulation. Neuromodulators, for instance, are known to modify effective neuronal excitabilities27. In our networks, gains and biases do not necessarily account for a single biophysical process but rather the coordinated effects of multiple processes. As long as the timescale of these processes is slower than the dynamics being predicted, we expect an approach similar to the one described here to be appropriate. However, this state-dependence may also imply that the inferred parameters do not generalize to new behavioral states.
Assessment of connectome-constrained solutions
It is known that recordings from different neurons are required to estimate neural dynamics lying in a manifold of linear dimensionality , independent of network size33. Connectome-constrained models go beyond such population-level descriptions of neural dynamics, as they are also concerned with how each specific neuron contributes to global activity patterns. This requires knowledge of unrecorded neurons’ loadings onto the low-dimensional manifold. This benchmark is appropriate when such models are used to predict the function of specific neurons or neuron types, or to guide experiments that manipulate specific neurons14;16;19.
The match between student and teacher activities in our models depends on multiple properties. These include whether random choices of single-neuron parameters produce similar dynamics in the two networks (Fig. 3 C, Extended Data Fig. 1), the extent of model mismatch (Fig. 4), and the degree to which training the student using activity recordings may compensate for the mismatch (Extended Data Fig. 1). These properties depend on specific features of the teacher network. Recent studies have demonstrated above-chance prediction of function using uniform or random parameters in models of the Drosophila nervous system14;19. In one case, accurate predictions of motor neuron responses to optogenetic stimulation was achieved without any adjustment of single-neuron parameters, which may reflect the presence of strong and direct excitatory pathways between stimulated neurons and output neurons14. In another case, further training of single-neuron parameters based on a task objective of detecting visual motion led to an improvement in predictions19. On the other hand, failure of a related approach in C. elegans was argued to be a consequence of model mismatch from unmodeled peptidergic interactions13.
We found that training student networks to match the teacher only led to improvements when, in addition to connectivity constraints, sufficiently many neurons’ activities were constrained. This result was independent of the initial performance of the system with random parameters (Extended Data Fig. 1). We focused on the improvement that can be achieved through knowledge provided by neural recordings, which was motivated by the observation that training the student only on the task-related readout of a teacher did not predict unrecorded neural activity (Fig. 1). However, it is possible that high-dimensional task readouts may improve predictions. Indeed, for our linear model, constraining activity along one task-related dimension is analogous to recording one additional neuron, as both correspond to a linear projection of the network’s vector of neural activities. In networks that perform multiple tasks or process diverse inputs, recording activity under multiple task or input conditions may improve the prediction of unrecorded activity, as has been argued for the Drosophila visual system19. This would likely require an assumption that single-neuron parameters are not strongly modulated across these conditions.
Properties of connectome-constrained and unconstrained network solutions
The loss functions of task-trained feedforward neural networks have been shown to exhibit multiple minima, with often counter-intuitive geometrical properties52;53. The multiplicity of minima arises from symmetries such as weight permutations in the network parameterization54. It remains unclear whether such ideas extend to recurrent neural networks. Our results for connectome-constrained networks are consistent with the existence of a single minumum or connected set of minima with stiff and sloppy parameter modes around the optimal solution (Fig. 7). The alignment between sloppy parameter modes in subsampled versus fully sampled loss functions explains the success in generalizing to unrecorded neurons.
Robustness to a large range of structural parameters and perturbations is a hallmark of biological systems, with a few stiff parameter combinations determining function45;55–59. We have shown that this is also true of connectome-constrained networks. One consequence of this observation is that, in underconstrained, data-driven models for neuroscience and machine learning, the distribution of parameters such as synaptic weights or single-neuron excitabilities found after successful training may not be predictive of task performance. Our work argues in favor of identifying stiff parameter combinations in such networks and using these to assess the similarity of network solutions60.
METHODS
Recurrent network models
We focused on recurrent neural networks where the activity of each neuron is described by a continuous variable, a firing rate , for . The firing rate of each neuron is calculated by applying a parametric function to the input that the neuron receives at each time point,
| (5) |
where denote the single-neuron parameters that modulate the function . We denote this input-to-rate function as the activation function. The activation function may depend on various single-neuron parameters , such as the gains and biases in Eq. (1).
The dynamics of the recurrent neural network follow
| (6) |
The matrix is the synaptic weight matrix, with each element indicating the signed synaptic strength of the connection from neuron to neuron . The constant indicates the timescale of single neurons. We assume the same for all neurons and set it equal to one, unless specified otherwise, such that all other timescales are given in units of . We denote the external input by , which includes the task-related input and can include additional private white noise provided to each neuron .
The dynamical landscape that a network can implement is thus determined by the order- single-neuron parameters and the synaptic weights. In the section Teacher RNNs below, we specify the choice of activation function, single-neuron parameters and synaptic weight matrices used in each figure.
Teacher-student: setup and training
We focus on a set of two RNNs. The first is the teacher RNN, which represents the network whose connectome is known and from which we can record neural activity. The other is the student RNN, which is trained (i.e., its parameters are optimized) to match the recorded neural activity in the teacher. Both networks’ dynamics are determined by Eq. (6). Asterisks are used to refer to teacher network parameters.
Unless otherwise specified, the teacher and the student network share the same synaptic weight matrix . Additionally, the external input and initial conditions are the same for teacher and student. The noise intensity is chosen to be zero in the teacher, since we focus on the case of no measurement noise. The noise intensity in the students is weak, to provide additional stability over training. The possible structural differences between teacher and student come from the set of single-neuron parameters . These single-neuron parameters are optimized so that student activity matches the recorded activity of the teacher. In Fig. 2 D and Extended Data Fig. 1, where we train students with unconstrained connectivity, we instead set the single-neuron parameters to be equal but the synaptic weight matrices to be different in teacher and student. In these cases, the student synaptic weights are trained.
The trained parameters are optimized to minimize the loss function
| (7) |
where the square brackets denote an average over the time points in the recorded window, and we sorted neurons such that the first neurons are those that are recorded. In Fig. 1, instead of defining the loss based on the recorded activity traces , we used the task readout, .
We trained the parameters of the student using standard gradient descent methods applied to time varying signals: we implemented backpropagation through time via the ADAM optimizer using Pytorch61–63. Learning rates varied between 0.0001 and 0.01, and decay rates of the first and second moments between 0.9 and 0.999.
Quantifying performance
We used two different metrics to assess the deviations between teacher and student. The first is the root mean-squared error. For Figs. 3 and 5, networks where the firing rates are very heterogeneous across neurons, or where the temporal profile of the responses is more relevant (which would be the case when comparing to, for instance, calcium fluorescence traces), we instead used a correlation-based score, which is not affected by the average amplitude of the responses. This score is defined as
| (8) |
where the angular brackets indicate the mean over neurons , and is the Pearson correlation coefficient for the -th neuron:
| (9) |
The square brackets indicate the average across time points, and is the time-averaged activity of the -th neuron. When there are multiple trials (Fig. 5 A–F), we calculated one score per neuron and trial, and then averaged over trials. For Fig. 3 C, we took the absolute value of before averaging over neurons, since we were interested in whether single neurons of the teacher reproduce the oscillatory dynamics of the teacher, independently of the sign, although our qualitative conclusions do not depend on this choice.
In Fig. 2, to match unrecorded neurons, we paired unrecorded neurons in the teacher with unrecorded neurons in the student at each training epoch, by solving the linear sum assignment problem using the function ‘linear_sum_assignment’ in Scipy. We calculated the mean squared error between all unrecorded neurons in the teacher, and unrecorded neurons in the student, and stored them in the matrix element . The linear sum assignment problem finds the matrix such that minimizes , with the constraint that each row in maps exactly one single neuron in the teacher to one neuron in the student.
Teacher RNNs and training parameters
In this section, we detail the choice of single-neuron parameters, the features of the teacher RNN, and the training parameters used in each figure. Unless otherwise specified, we used in units of the single-neuron time constant. We injected noise at each timestep of the dynamics with standard deviation 0.002. See also the shared code (to be made available upon publication), for reproducing all the numerical experiments in the study.
Figures 1 and 2: The teacher RNN () is trained with learning rate 0.001 during 1400 epochs to solve the cycling task. The initial connectivity is chosen as follows: first, each synapse is drawn from a random Gaussian distribution with mean zero and variance . The sparse weights (fraction ) are randomly chosen, set to zero, and not trained. A fraction of neurons selected randomly is set to be excitatory, such that their synaptic strengths are set to their absolute value, while the remaining fraction of neurons, , is set to be inhibitory. Synapses are rectified to their assigned sign after each training epoch. Trials are 20 time units long. The single-neuron activation function is given by Softplus:
| (10) |
where we set the smoothness parameter to .
The student RNNs with unknown single-neuron parameters are trained for 7000 epochs and learning rate 0.001. The student RNN with unconstrained connectivity shares the same single-neuron parameters as the teacher RNN, to facilitate the comparison. Fig. 2 shows results for 14 different trained students. The learning rate was set to 0.005. The synaptic weights are initialized randomly following a Gaussian distribution, and the weight signs are correctly assigned (i.e., the student knows whether a neuron is excitatory or inhibitory). To compare the weights after training, we picked a random unrecorded neuron from the teacher and matched it with the unrecorded neuron with the most similar activity profile. Then, the selected neurons in teacher and student are discarded, and we picked a new neuron to be matched in the teacher. This procedure is repeated until all neurons are paired.
Figure 3: The teacher networks in the top row are rank-two networks whose synaptic weight matrices are given by:
| (11) |
The activation function for each neuron is the tanh function, and we consider gains as the only single-neuron parameter. Networks have variable numbers of neurons , but the distribution of connectivity loadings is given by fixed parameters, such that all the networks generate the same dynamics for large . The connectivity loadings of the -th neuron, are sampled from a four-variable Gaussian distribution with zero mean. The variance parameters are for , and . These parameters lead to a limit cycle in the dynamics36;47.
The gains are chosen to be Gaussian, with mean 1 and standard deviation 0.9, uncorrelated with all the other connectivity loadings. The precise shape of the gain distribution does not affect the dynamics, only its mean and variance, as long as the gains are uncorrelated with the connectivity loadings. We selected trajectories that start on the limit cycle, and evolve during 20 time units. The student network is initialized in this case with homogeneous unitary gains, and is trained for 7000 epochs with learning rate 0.005.
The teacher networks in the bottom row have random synaptic weights as in37, where the are randomly drawn from a Gaussian distribution with mean zero and standard deviation . The single-neuron gains are drawn from a Gaussian distribution of unit mean and standard deviation 0.5. One trial with fixed and known initial conditions is considered.
The student networks were initialized before training with homogeneous gains, .
Figure 4: The teacher network is the same biologically-inspired network as in Fig. 2.
Figure 5: Drosophila larva (top row). The teacher network was trained, following Zarin et al. (2019)16 such that the motor neurons produce activation sequences consistent with forward and backward crawling, with an additional L2-regularization loss on the activity of premotor units. The synaptic weights were fixed based on existing synapses, using neurotransmitter identity when available, and normalizing based on the percent input received by the postsynaptic target, as in Zarin et al. (2019)16. There is a different tonic input for each motion type. The trained parameters are the biases, gains and the tonic input patterns for the two motion directions. The single-neuron time constant in premotor and motor neurons is 0.2, the activation function is Softplus. Gains were bounded during training to be between 0.5 and 5.
The student network focused on the 178 premotor neurons, during both forward and backward motion. Gains and biases were trained starting from a random permutation of the teacher’s parameters. We added a small L2-regularization penalty to the loss function to reduce instabilities in the solutions.
Drosophila adult (central complex). The orientation selectivity of EPG neurons is assigned based on Turner-Evans et al. (2021)40, where each EPG cell type is maximally selectively 22.5° from their neighboring EPGs, tiling the whole range of angular directions. The teacher was trained such that the EPG neurons are able to produce a bump of activity for 2.5 time units, 2 time units after a brief pulse (0.3 time units) is presented. The activities of PEN1s, PEN2s, PEG and Delta7s were not constrained. The nonlinearity was Softplus, with a parameter of . The signs of the synaptic weights, based on the hemibrain connectome, were given by the predicted neurotransmitter, and the strength of each connection was set proportional to the number of synapses. The overall scaling of the synaptic weight matrix was set such that the largest eigenvalue has real part of 0.8, and then gains and biases were trained.
To build the student network, the gains and biases of each cell type were shuffled within their own cell group. The gains were further scaled down by a factor 0.8. The gains were forced to be non-negative. We used 60 different trials, with randomly set orientations.
Larval zebrafish (hindbrain). The teacher network is a linear network based on Vishwanathan (2024)41. The network was not trained, the overall strength of recurrent connections was fixed such that the real part of the largest eigenvalue is less but close to zero. The input pattern was randomly set, but made sure to overlap with the subspace of the slowest activity mode. We then randomly assigned a set of gains from a lognormal distribution with mean 1 and standard deviation 0.3 to each neuron. In order to keep the same dynamics, we rescaled each column of the synaptic weight matrix by the inverse of the corresponding gain.
Figure 6: For the connectivity of the teacher in Fig. 6 B–H, we drew each synaptic strength from Gaussian distribution with mean zero and standard deviation , and based on the singular value decomposition, kept the first 60 rank-one components. The network size was .
In Fig. 6 F (inset), to calculate the principle angle between the first singular vectors of the subsampled matrix and the full matrix , we calculated the left singular vectors of the matrix , and the first left singular vectors of the matrix . The principle angle measures the maximum angle between two linear subspaces. We computed the principal angle as the maximum singular value of the matrix product , for .
Figure 7: We designed a teacher RNN with a rank-two synaptic weight matrix and two populations36, tanh activation function and gains as single-neuron parameters, the biases are set to zero, neurons. The parameters in the first population are: , , and ; and in the second population , and . The gains are reduced to a two-dimensional parameter space, where all the gains of neurons in population 1 have the same value, , and all the gains of neurons in population 2 have value . In the teacher network, and . The parameters are chosen such that the first population has more control of the dynamics along the variable , and the second population controls .
Figure 8: The linear network corresponds to the same network as in Fig. 6. The nonlinear network is the same network as in Fig. 2.
Statistics & Reproducibility
For each teacher network, we trained at least 10 student networks in order to obtain robust estimates of the prediction accuracy. All trained student networks were included in the analysis. Student networks whose activity diverged during training were retrained with a different random seed until convergence. No other data were excluded from the analyses.
Prediction in linear recurrent networks
In the linear model, the RNNs are linear networks with dynamics
| (12) |
We define the activity in this linear network as . We assume that the real part of all eigenvalues of the connectivity matrix are smaller than unity, so that the linear dynamics are stable. The single-neuron parameters correspond to the bias. Throughout the results section, we focused on the fixed point activity, which is given in vector form by
| (13) |
where is the identity matrix. There is a linear mapping between single-neuron parameters and activity , given my a matrix , in this case defined as . The notation indicates the pseudo-inverse.
In linear networks, such as the simplified model studied here, it is possible to formulate the teacher-student setup as a system identification problem, and estimate the parameters using alternative approaches to gradient-based training, such as subspace methods64. For consistency with our approach to nonlinear networks, we use gradient descent to optimize the parameters of the student. However, the same conclusions can be reached from a system identification perspective (see Supplementary Note).
Fully sampled teacher
The loss function when all neurons are recorded, is given by the quadratic form
| (14) |
such that there is one global minimum when the student and teacher are identical to each other, , and the Hessian of the loss is independent of the teacher’s biases . Running gradient descent, in the limit of small learning rates , leads to Eq. (4) for the estimated biases in the student over the timecourse of learning, which reads in vector form:
| (15) |
Equation (15) shows that the evolution of single-neuron parameters is given by a linear dynamical system. The eigenvalue decomposition of , or equivalently, the singular vector decomposition of , therefore determines how fast and along which modes the parameters decay toward the teacher values, . Given the singular vector decomposition , we denote the left singular vector an activity mode, and right singular vector a parameter mode. The error in parameter mode decreases over training with timescale , reducing the error in activity mode . An initial guess which is a distance of 1 away from the teacher along mode generates an error in neural activity along mode of magnitude . Thus, parameter modes that have large effects on activity are learned quickly, while parameter modes that have small effects on activity are learned more slowly. We refer to parameter modes corresponding to large and small singular values as stiff and sloppy modes, respectively.
If the connectivity is not full-rank, some singular values of the mapping matrix will be zero. In that case, the parameter values along the modes corresponding to singular value (the extreme case of sloppy parameter modes) cannot be inferred through gradient descent, although that mismatch does not cause any error in the activity of unrecorded neurons.
All the results can be directly extended to linear networks where transient trajectories are considered, given an initial state . For time-dependent responses, the dynamics follow
| (16) |
where there is an affine mapping from to the parameters , given by :
| (17) |
The second term in Eq. (16) is the same for the teacher and the student because we assume that the initial state is known. Thus, the difference in activity between networks is:
| (18) |
The loss function is the time-averaged squared error of the activity:
| (19) |
where the square brackets indicate a time average. The matrix that determines the stiff and sloppy modes is therefore the time-averaged matrix . The eigenvalues of this matrix determine the level of stiffness, and the eigenvectors determine the parameter modes.
Note that in this model, we have assumed that all the recurrent dimensions are explored by the teacher, such that the rank of the connectivity determines the dimensionality of the activity. In practice, neural activity is recorded for a limited time window in response to a small set of inputs, so the dimensionality of the activity is much lower. The rank of the connectivity sets an upper bound on the network’s activity (see Extended Data Fig. 4 for a comparison of the dimensionality of activity and rank of the connectivity in connectome-constrained recurrent networks).
Subsampled activity
Recording from a subset of neurons is equivalent to selecting the rows of matrix corresponding to the recorded neurons and removing the rest. We refer to this matrix as matrix . Equations (14) and (15) still hold, when substituting for .
The effect of subsampling limits the number of learnable or stiff parameter modes of the loss function used for training, which cannot exceed . The fact that the initial parameter guess can only be corrected along modes makes the error in unrecorded activity non-zero when the rank of is larger than , i.e. when not enough neurons are sampled. Furthermore, the parameter modes and activity modes without a non-zero eigenvalue of the training loss need not align with the stiffest modes of the fully sampled loss function.
One recorded neuron
In linear networks, we can calculate the average error when we record only from neuron . We use the vector to refer to the row of the subsampled matrix . After training has converged for a student with initial parameters , which is equivalent to assuming zero training error for the recorded activity of neuron (given the absence of measurement noise), the vector of biases after training, is
| (20) |
The squared error in single-neuron parameters (combining both recorded and unrecorded neurons), , is calculated based on the norm of the vector given by Eq. 20, which reads:
| (21) |
Assuming that the initial guesses are unbiased with respect to the teacher parameters , on average over initial conditions, the improvement in the error in parameters is
| (22) |
Therefore, on average, the error in parameter space is equally reduced for any selected neuron.
Similarly, the error in the activity of neurons reads:
| (23) |
such that the squared error, using the singular value decomposition of , is
| (24) |
The squared error can be larger or smaller than the error before training, unlike the error in parameter space, which can only decrease. Nevertheless, on average over initial conditions, the expected error always decreases and is given by
| (25) |
where is the angle between and . Eq. 25 is used to calculate the theoretical predictions in Fig. 8 E.
Error in unrecorded neural activity vs. single-neuron parameters
From the perspective of a single neuron , we can write the following identity using the equation for the linear network dynamics at the fixed point:
| (26) |
The first term corresponds to the error due to incorrect prediction of the activity of other neurons in the network, while the second term corresponds to the error due to parameter mismatch between teacher and student.
If neuron is a recorded neuron, then after training has converged, Eq. (26) equals zero, imposing the constraint
| (27) |
In other words, the weighted sum of errors from incorrectly inferring parameters (r.h.s.) compensates for the weighted sum of errors from the incorrect prediction of activity (l.h.s.).
If neuron is not a recorded neuron, then both terms in Eq. (26) in general contribute to the squared error. Which term has a stronger contribution depends on the strength of recurrent connectivity. For strong recurrence, the first term will dominate, while for weakly connected networks, the second term will dominate. As more neurons are recorded, the amplitudes of both contributions decay similarly.
Optimal selection of neurons: linear RNN
To calculate the best and worst strategy for sampling neurons (Fig. 8) we used a greedy strategy, where we first selected the neuron with the highest and lowest expected reduction in activity error, based on Eq. 25. Then, we proceeded iteratively, projecting out the component from the -th selected neuron from the matrix , calculating the matrix :
| (28) |
We then selected again the row-vector that maximizes (minimizes) the decrease error in Eq. 25, for the best (worst) greedy selection of neurons.
Optimal selection of neurons: nonlinear RNN
For any teacher RNN with unknown gains or nonlinear activation functions, the mapping between unknown single-neuron parameters and activity is not given by a linear transformation via a matrix . Moreover, the linearization of the gradient dynamics (Eq. 15) close to the teacher parameter depends on the specific parameters, unlike the linear case. Nevertheless, we can still compute the best and worst selection of neurons based on an initial guess of the target parameters.
We focus on networks with firing rates given by , where the notation indicates a diagonal matrix whose non-zero elements are given by vector , and we assume the function is invertible. We are interested in the linearization and . We are focused on fixed point activity and thus, using Eq. (6), we can define the function:
| (29) |
By applying the implicit function theorem to , we can calculate the linearized mapping from parameters to activity:
| (30) |
| (31) |
This linear relationship is analogous to the parameter-to-activity mapping defined previously for linear RNNs, allowing us to use the same procedure iteratively. This amounts to assuming that the curvature of the loss function close to the current parameter estimate is similar to the curvature close to the teacher parameters.
In Fig. 7 G–H, the Jacobian of the mapping between time-varying activity and single-neuron parameters around the teacher’s parameter values was estimated numerically, using Pytorch’s automatic differentiation.
Extended Data
Extended Data Figure 1: Related to Fig. 2.

A Teacher as in Fig. 1. The students are trained on a varying number of recorded neurons M.
B Average error in the recorded and unrecorded activity between teacher and students.
C Left: Error in the network activity for a given student network in a given trial, when M = 20 neurons are recorded. Right: Error in the task-related readout signal. While the recorded neurons have low error, the unrecorded neurons in the student display large deviations.
D Analogous to C, when more neurons are recorded, M = 60. In this case, the activity of unrecorded neurons and the readout are well predicted.
E Teacher network from panel A receives a strong external two-dimensional time-varying input, fed to a subset of 100 excitatory neurons. Middle: The dimensionality of the activity, measured by the participation ratio, increases with the input.
F Error in unrecorded neurons after training student networks to match the input-driven teacher (color dots), compared to the non-driven teacher (grey dots). Fewer recorded neurons are required to predict activity of unrecorded neurons in this example input-driven network.
G Input-driven teacher network with different levels of connectivity sparsity and gain heterogeneity. Teachers have E-I random connectivity, and are initialized at the fixed point. A positive input of unit strength is delivered to 5 excitatory neurons. Recorded neurons correspond to excitatory neurons, while unrecorded neurons can be both excitatory or inhibitory. Teacher networks are generated with different fractions f of non-zero weights, and different ranges for the uniformly distributed gains. Both gains and biases are trained in the students.
H Error in unrecorded activity after training vs number of recorded neurons, for different level of sparsity f and gain distributions. While the overall magnitude of the error changes for different gain strengths, the decay of the error as a function of M does not change.
Extended Data Figure 2: Related to Fig. 3. Teacher networks with different dynamics.

A Teachers with variable network size and fixed rank-two connectivity, generating a limit cycle. Right: Error in the activity of recorded neurons after training. The students always learn the dynamics of the teacher.
B Error in the single-neuron gains after training.
C Example of error in the activity of a recorded neuron and an unrecorded neuron, when there is only one recorded neuron (left), compared to when 7 neurons are recorded (right). For one recorded neuron, the student learns the frequency of the limit cycle, but the temporal profile of the unrecorded neurons does not much the profile of the teacher network. Example for N = 400.
D Teachers with variable network size and random connectivity, generating chaotic dynamics. Right: Error in the activity of recorded neurons after training. The students always learn the dynamics of the teacher.
E Error in the single-neuron gains for the chaotic teachers. Note that the single-neuron parameters are much better inferred given enough recorded neurons when the teacher is chaotic than when it is low-rank, because there are many more stiff dimensions.
F Traces of one example neuron in teacher and student networks with size N = 400 (left) and N = 1000 (right). For N = 400, M = 64 recorded neurons is sufficient to accurately match unrecorded neural activity from the teacher (gray line), while for N = 1000, M = 64 recorded neurons is insufficient but M = 256 is sufficient.
Extended Data Figure 3: Related to Fig. 4. Training connectivity with model mismatch between teacher and student.

A Teacher with model mismatch in the activation function, from Fig. 4 A–C.
B Example traces of one recorded neuron and one unrecorded neuron in the teacher and after training the student with mismatch in the β parameter. The students networks were trained with 20 recorded neurons (left) and with 150 recorded neurons (right).
C Teacher-student framework with mismatch. We train the connectivity of the student, given the teacher’s connectivity as initial condition. The single-neuron parameters are the same in teacher and student, while there is a mismatch in the activation function. Same network as in Fig. 4.
D The activation function is a smooth rectification but with different degrees of smoothness, parameterized by a parameter β. Teacher RNN from Fig. 2.
E Errors in the activity of recorded (left) and unrecorded (right) neurons for different values of model mismatch between teacher and student. We observe a minor decrease in the error in unrecorded neurons when recording from a large number of neurons, M ≈ 150.
F Error in the recorded activity (loss function) for three different mismatch values as a function of training epochs (β = 1. means no mismatch).
G Error in the unrecorded activity (loss function) for three different mismatch values as a function of training epochs.
H Removing the mismatch in activation by training an additional parameter. We train a student network with the same connectivity as the teacher and different single-neuron parameters. However, the student also does not know the smoothness parameter β. The trained parameters are therefore the gains and biases of each neuron and the smoothness β.
I Error in unrecorded activity after training on a subset of M recorded units, similar to C. Training the smoothness parameter of the nonlinearity provides the student with the same prediction power as students without mismatch (see Fig. 2).
J Estimated parameter β during training (average and SEM over 10 different initializations). Networks do not retrieve the exact teacher value (β* = 1) although converge to values not far from it on average. Students have a bias towards estimating sharper activation functions (β > 1). Both bias and variance get reduced as the number of recorded neurons is increased.
Extended Data Figure 4: Related to Fig. 5. Dimensionality of the activity and rank of connectivity in the data-constrained RNNs.

A Neural activity traces (centered) used for training the student networks for the three different data constrained RNNs: the premotor network in the Drosophila larva, the central complex in the adult Drosophila, and the oculomotor integrator in larval zebrafish. Different trials/conditions have been concatenated.
B Left: First eigenvalues of the covariance spectrum of the datasets. Right: Participation ratio of the activity covariance. The dimensionality of neural activity is higher in the premotor system, then the CX and then the premotor network, indicated by how fast the eigenvalues decay.
C Left: Singular values of the connectivity matrix. Right: Estimated rank of the connectivity matrix J, calculated using the participation ratio of the distribution of singular values of J. Given the sparsity and heterogeneity in connectomes, the rank of the connectivity is high.
Supplementary Material
ACKNOWLEDGEMENTS
The authors are grateful to L.F. Abbott for helpful discussions and comments on the manuscript. M.B. and A.L.-K were supported by the Kavli Foundation, the Gatsby Charitable Foundation GAT3708, the Burroughs Wellcome Foundation, and NIH awards R01EB029858 and RF1DA060772. A.L.-K. was supported by the McKnight Endowment Fund. The funders had no role in study design, data collection and analysis, decision to publish or preparation of the manuscript.
Footnotes
COMPETING INTERESTS
The authors declare no competing interests.
CODE AVAILABILITY
All the simulations and analyses were performed using custom code written in Python (https://www.python.org). The code used to generate all the results and can be found in ref.65 and https://github.com/emebeiran/connconstr.
DATA AVAILABILITY
The connectomics data used in this study were published in Zarin et al.16 for Drosophila larva, in Scheffer et al.9 for the central complex of adult Drosophila, and in Vishwanathan et al.41 for the brainstem of the larval zebrafish. All the generated data shown in the main results, together with the teacher and student recurrent networks is publicly available on https://doi.org/10.5281/zenodo.1661835365.
REFERENCES
- [1].Das A, Fiete IR, Systematic errors in connectivity inferred from activity in strongly recurrent networks, Nature Neuroscience 2020 23:10 23 (10) (2020) 1286–1296. [Google Scholar]
- [2].Haber A, Schneidman E, Learning the architectural features that predict functional similarity of neural networks, Physical Review X 12 (2) (2022) 021051. [Google Scholar]
- [3].Levina A, Priesemann V, Zierenberg J, Tackling the subsampling problem to infer collective properties from limited data, Nature Reviews Physics 2022 4:12 4 (12) (2022) 770–784. [Google Scholar]
- [4].Liang T, Brinkman BA, Statistically inferred neuronal connections in subsampled neural networks strongly correlate with spike train covariances, Physical Review E 109 (4) (2024) 044404. [DOI] [PubMed] [Google Scholar]
- [5].Dinc F, Shai A, Schnitzer M, Tanaka H, Cornn: Convex optimization of recurrent neural networks for rapid inference of neural dynamics, Advances in Neural Information Processing Systems 36 (2023) 51273–51301. [Google Scholar]
- [6].White JG, Southgate E, Thomson JN, Brenner S, The Structure of the Nervous System of the Nematode Caenorhabditis elegans, Philosophical Transactions of the Royal Society of London. Series B, Biological Sciences 314 (1165) (1986) 1–340. [DOI] [PubMed] [Google Scholar]
- [7].Ohyama T, Schneider-Mizell CM, Fetter RD, Aleman JV, Franconville R, Rivera-Alba M, Mensh BD, Branson KM, Simpson JH, Truman JW, Cardona A, Zlatic M, A multilevel multimodal circuit enhances action selection in Drosophila, Nature 520 (7549) (2015) 633–639. [DOI] [PubMed] [Google Scholar]
- [8].Zheng Z, Lauritzen JS, Perlman E, Saalfeld S, Fetter RD, Bock Correspondence DD, A Complete Electron Microscopy Volume of the Brain of Adult Drosophila melanogaster, Cell 174 (2018) 730–743.e22. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [9].Scheffer LK, Xu CS, Januszewski M, Lu Z, Takemura SY, Hayworth KJ, Huang GB, Shinomiya K, Maitin-Shepard J, Berg S, Clements J, Hubbard PM, Katz WT, Umayam L, Zhao T, Ackerman D, Blakely T, Bogovic J, Dolafi T, Kainmueller D, Kawase T, Khairy KA, Leavitt L, Li PH, Lindsey L, Neubarth N, Olbris DJ, Otsuna H, Trautman ET, Ito M, Bates AS, Goldammer J, Wolff T, Svirskas R, Schlegel P, Neace ER, Knecht CJ, Alvarado CX, Bailey DA, Ballinger S, Borycz JA, Canino BS, Cheatham N, Cook M, Dreher M, Duclos O, Eubanks B, Fairbanks K, Finley S, Forknall N, Francis A, Hopkins GP, Joyce EM, Kim S, Kirk NA, Kovalyak J, Lauchie SA, Lohff A, Maldonado C, Manley EA, McLin S, Mooney C, Ndama M, Ogundeyi O, Okeoma N, Ordish C, Padilla N, Patrick C, Paterson T, Phillips EE, Phillips EM, Rampally N, Ribeiro C, Robertson MK, Rymer JT, Ryan SM, Sammons M, Scott AK, Scott AL, Shinomiya A, Smith C, Smith K, Smith NL, Sobeski MA, Suleiman A, Swift J, Takemura S, Talebi I, Tarnogorska D, Tenshaw E, Tokhi T, Walsh JJ, Yang T, Horne JA, Li F, Parekh R, Rivlin PK, Jayaraman V, Costa M, Jefferis GS, Ito K, Saalfeld S, George R, Meinertzhagen IA, Rubin GM, Hess HF, Jain V, Plaza SM, A connectome and analysis of the adult Drosophila central brain, eLife 9 (2020) 1–74. [Google Scholar]
- [10].Dorkenwald S, Matsliah A, Sterling AR, Schlegel P, Yu S-C, McKellar CE, Lin A, Costa M, Eichler K, Yin Y, et al. , Neuronal wiring diagram of an adult brain, Nature 634 (8032) (2024) 124–138. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [11].Hildebrand DGC, Cicconet M, Torres RM, Choi W, Quan TM, Moon J, Wetzel AW, Scott Champion A, Graham BJ, Randlett O, Plummer GS, Portugues R, Bianco IH, Saalfeld S, Baden AD, Lillaney K, Burns R, Vogelstein JT, Schier AF, Lee WCA, Jeong WK, Lichtman JW, Engert F, Whole-brain serial-section electron microscopy in larval zebrafish, Nature 545 (7654) (2017) 345–349. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [12].Turner MH, Mann K, Clandinin TR, The connectome predicts resting-state functional connectivity across the Drosophila brain, Current biology 31 (11) (2021) 2386–2394.e3. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [13].Randi F, Sharma AK, Dvali S, Leifer AM, Neural signal propagation atlas of Caenorhabditis elegans, Nature 2023 623:7986 623 (7986) (2023) 406–414. [Google Scholar]
- [14].Shiu PK, Sterne GR, Spiller N, Franconville R, Sandoval A, Zhou J, Simha N, Kang CH, Yu S, Kim JS, et al. , A Drosophila computational brain model reveals sensorimotor processing, Nature 634 (8032) (2024) 210–219. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [15].Pospisil DA, Aragon MJ, Dorkenwald S, Matsliah A, Sterling AR, Schlegel P, Yu S.-c., McKellar CE, Costa M, Eichler K, et al. , The fly connectome reveals a path to the effectome, Nature 634 (8032) (2024) 201–209. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [16].Zarin AA, Mark B, Cardona A, Litwin-Kumar A, Doe CQ, A multilayer circuit architecture for the generation of distinct locomotor behaviors in Drosophila, Elife 8 (2019) e51781. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [17].Keller AJ, Dipoppa M, Roth MM, Caudill MS, Ingrosso A, Miller KD, Scanziani M, A Disinhibitory Circuit for Contextual Modulation in Primary Visual Cortex, Neuron 108 (6) (2020) 1181–1193.e8. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [18].Kohn JR, Portes JP, Christenson MP, Abbott LF, Behnia R, Flexible filtering by neural inputs supports motion computation across states and stimuli, Current biology 31 (23) (2021) 5249–5260.e5. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [19].Lappalainen JK, Tschopp FD, Prakhya S, McGill M, Nern A, Shinomiya K, Takemura S.-y., Gruntman E, Macke JH, Turaga SC, Connectome-constrained networks predict neural activity across the fly visual system, Nature 634 (8036) (2024) 1132–1140. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [20].Eckstein N, Bates AS, Champion A, Du M, Yin Y, Schlegel P, Lu AK-Y, Rymer T, Finley-May S, Paterson T, et al. , Neurotransmitter classification from electron microscopy images at synaptic sites in Drosophila melanogaster, Cell 187 (10) (2024) 2574–2594. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [21].Barnes CL, Bonnéry D, Cardona A, Synaptic counts approximate synaptic contact area in Drosophila, PLOS ONE 17 (4) (2022) e0266064. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [22].Kasai H, Fukuda M, Watanabe S, Hayashi-Takagi A, Noguchi J, Structural dynamics of dendritic spines in memory and cognition, Trends in neurosciences 33 (3) (2010) 121–129. [DOI] [PubMed] [Google Scholar]
- [23].Bargmann CI, Marder E, From the connectome to brain function, Nature methods 10 (6) (2013) 483–490. [DOI] [PubMed] [Google Scholar]
- [24].Marder E, Neuromodulation of neuronal circuits: back to the future, Neuron 76 (1) (2012) 1–11. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [25].Gutierrez GJ, O’Leary T, Marder E, Multiple mechanisms switch an electrically coupled, synaptically inhibited neuron between competing rhythmic oscillators, Neuron 77 (5) (2013) 845–858. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [26].Stroud JP, Porter MA, Hennequin G, Vogels TP, Motor primitives in space and time via targeted gain modulation in cortical networks, Nature Neuroscience 2018 21:12 21 (12) (2018) 1774–1783. [Google Scholar]
- [27].Ferguson KA, Cardin JA, Mechanisms underlying gain modulation in the cortex, Nature Reviews Neuroscience 2020 21:2 21 (2) (2020) 80–92. [Google Scholar]
- [28].Connors BW, Gutnick MJ, Intrinsic firing patterns of diverse neocortical neurons, Trends in neurosciences 13 (3) (1990) 99–104. [DOI] [PubMed] [Google Scholar]
- [29].Kubota Y, Hatada S, Kondo S, Karube F, Kawaguchi Y, Neocortical inhibitory terminals innervate dendritic spines targeted by thalamocortical afferents, The Journal of neuroscience : the official journal of the Society for Neuroscience 27 (5) (2007) 1139–1150. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [30].Perich MG, Arlt C, Soares S, Young ME, Mosher CP, Minxha J, Carter E, Rutishauser U, Rudebeck PH, Harvey CD, Rajan K, Inferring brain-wide interactions using data-constrained recurrent neural network models, bioRxiv (2021) 2020.12.18.423348. [Google Scholar]
- [31].Seung HS, Sompolinsky H, Tishby N, Statistical mechanics of learning from examples, Physical Review A 45 (8) (1992) 6056. [Google Scholar]
- [32].Saad D, Solla SA, Exact Solution for On-Line Learning in Multilayer Neural Networks, Physical Review Letters 74 (21) (1995) 4337. [DOI] [PubMed] [Google Scholar]
- [33].Gao P, Trautmann E, Yu B, Santhanam G, Ryu S, Shenoy K, Ganguli S, A theory of multineuronal dimensionality, dynamics and measurement, bioRxiv (2017) 214262. [Google Scholar]
- [34].Kim CM, Finkelstein A, Chow CC, Svoboda K, Darshan R, Distributing task-related neural activity across a cortical network through task-independent connections, Nature Communications 2023 14:1 14 (1) (2023) 1–21. [Google Scholar]
- [35].Russo AA, Bittner SR, Perkins SM, Seely JS, London BM, Lara AH, Miri A, Marshall NJ, Kohn A, Jessell TM, Abbott LF, Cunningham JP, Churchland MM, Motor cortex embeds muscle-like commands in an untangled population response, Neuron 97 (4) (2018) 953. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [36].Beiran M, Dubreuil A, Valente A, Mastrogiuseppe F, Ostojic S, Shaping Dynamics With Multiple Populations in Low-Rank Recurrent Networks, Neural Computation 33 (6) (2021) 1572–1615. [DOI] [PubMed] [Google Scholar]
- [37].Sompolinsky H, Crisanti A, Sommers HJ, Chaos in Random Neural Networks, Physical Review Letters 61 (3) (1988) 259. [DOI] [PubMed] [Google Scholar]
- [38].Clark DG, Abbott LF, Litwin-Kumar A, Dimension of Activity in Random Neural Networks, Physical Review Letters 131 (11) (2023) 118401. [DOI] [PubMed] [Google Scholar]
- [39].Hamood AW, Marder E, Animal-to-Animal Variability in Neuromodulation and Circuit Function, Cold Spring Harbor symposia on quantitative biology 79 (2014) 21–28. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [40].Turner-Evans DB, Jensen KT, Ali S, Paterson T, Sheridan A, Ray RP, Wolff T, Lauritzen JS, Rubin GM, Bock DD, et al. , The neuroanatomical ultrastructure and function of a biological ring attractor, Neuron 109 (9) (2021) 1582. [DOI] [PubMed] [Google Scholar]
- [41].Vishwanathan A, Sood A, Wu J, Ramirez AD, Yang R, Kemnitz N, Ih D, Turner N, Lee K, Tartavull I, et al. , Predicting modular functions and neural coding of behavior from a synaptic wiring diagram, Nature Neuroscience (2024) 1–12. [Google Scholar]
- [42].Kim SS, Rouault H, Druckmann S, Jayaraman V, Ring attractor dynamics in the Drosophila central brain, Science 356 (6340) (2017) 849–853. [DOI] [PubMed] [Google Scholar]
- [43].Noorman M, Hulse BK, Jayaraman V, Romani S, Hermundstad AM, Maintaining and updating accurate internal representations of continuous variables with a handful of neurons, Nature Neuroscience (2024) 1–11. [Google Scholar]
- [44].Ben-Yishai R, Bar-Or RL, Sompolinsky H, Theory of orientation tuning in visual cortex., Proceedings of the National Academy of Sciences 92 (9) (1995) 3844–3848. [Google Scholar]
- [45].Prinz AA, Bucher D, Marder E, Similar network activity from disparate circuit parameters, Nature Neuroscience 7 (12) (2004) 1345–1352. [DOI] [PubMed] [Google Scholar]
- [46].Mastrogiuseppe F, Ostojic S, Linking Connectivity, Dynamics, and Computations in Low-Rank Recurrent Neural Networks, Neuron 99 (3) (2018) 609–623.e29. [DOI] [PubMed] [Google Scholar]
- [47].Dubreuil A, Valente A, Beiran M, Mastrogiuseppe F, Ostojic S, The role of population structure in computations through neural dynamics, Nature neuroscience 25 (6) (2022) 783–794. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [48].Singer W, Synchronization of cortical activity and its putative role in information processing and learning, Annual review of physiology 55 (1) (1993) 349–374. [Google Scholar]
- [49].Koch C, Segev I, The role of single neurons in information processing, Nature neuroscience 3 (11) (2000) 1171–1177. [DOI] [PubMed] [Google Scholar]
- [50].Goaillard J-M, Marder E, Ion channel degeneracy, variability, and covariation in neuron and circuit resilience, Annual review of neuroscience 44 (2021) 335–357. [Google Scholar]
- [51].Seung HS, Predicting visual function by interpreting a neuronal wiring diagram, Nature 634 (8032) (2024) 113–123. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [52].Li H, Xu Z, Taylor G, Studer C, Goldstein T, Visualizing the Loss Landscape of Neural Nets, Advances in Neural Information Processing Systems 31 (2018). [Google Scholar]
- [53].Fort S, Jastrzebski S, Large Scale Structure of Neural Network Loss Landscapes, Advances in Neural Information Processing Systems 32 (2019). [Google Scholar]
- [54].Simsek B, Ged F, Jacot A, Spadaro F, Hongler C, Gerstner W, Brea J, Geometry of the Loss Landscape in Overparameterized Neural Networks: Symmetries and Invariances (jul 2021).
- [55].Gutenkunst RN, Waterfall JJ, Casey FP, Brown KS, Myers CR, Sethna JP, Universally Sloppy Parameter Sensitivities in Systems Biology Models, PLOS Computational Biology 3 (10) (2007) e189. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [56].Daniels BC, Chen YJ, Sethna JP, Gutenkunst RN, Myers CR, Sloppiness, robustness, and evolvability in systems biology, Current opinion in biotechnology 19 (4) (2008) 389–395. [DOI] [PubMed] [Google Scholar]
- [57].Fisher D, Olasagasti I, Tank DW, Aksay ER, Goldman MS, A modeling framework for deriving the structural and functional architecture of a short-term memory microcircuit, Neuron 79 (5) (2013) 987–1000. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [58].Naumann EA, Fitzgerald JE, Dunn TW, Rihel J, Sompolinsky H, Engert F, From whole-brain data to functional circuit models: the zebrafish optomotor response, Cell 167 (4) (2016) 947–960. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [59].Otopalik AG, Goeritz ML, Sutton AC, Brookings T, Guerini C, Marder E, Sloppy morphological tuning in identified neurons of the crustacean stomatogastric ganglion, eLife 6 (feb 2017). [Google Scholar]
- [60].O’Leary T, Sutton AC, Marder E, Computational models in the age of large datasets, Current opinion in neurobiology 32 (2015) 87–94. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [61].Werbos PJ, Backpropagation Through Time: What It Does and How to Do It, Proceedings of the IEEE 78 (10) (1990) 1550–1560. [Google Scholar]
- [62].Kingma DP, Ba JL, Adam: A method for stochastic optimization, in: 3rd International Conference on Learning Representations - Conference Track Proceedings, International Conference on Learning Representations, ICLR, 2015. [Google Scholar]
- [63].Paszke A, Gross S, Chintala S, Chanan G, Yang E, Facebook ZD, Research AI, Lin Z, Desmaison A, Antiga L, Srl O, Lerer A, Automatic differentiation in PyTorch, in: Advances in Neural Information Processing Systems, 2017, pp. 8024–8035. [Google Scholar]
- [64].Van Overschee d P., De Moor B, Subspace identification for linear systems: Theory—Implementation—Applications, Springer Science & Business Media, 2012. [Google Scholar]
- [65].Beiran M, Litwin-Kumar A, Dataset and code for generating the figures of publication Prediction of neural activity in connectome- constrained recurrent networks, Zenodo, 10.5281/zenodo.16618353 (2025) [DOI] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Data Availability Statement
The connectomics data used in this study were published in Zarin et al.16 for Drosophila larva, in Scheffer et al.9 for the central complex of adult Drosophila, and in Vishwanathan et al.41 for the brainstem of the larval zebrafish. All the generated data shown in the main results, together with the teacher and student recurrent networks is publicly available on https://doi.org/10.5281/zenodo.1661835365.
