Skip to main content
Proceedings of the National Academy of Sciences of the United States of America logoLink to Proceedings of the National Academy of Sciences of the United States of America
. 2026 Jan 9;123(2):e2426883123. doi: 10.1073/pnas.2426883123

Primate-informed neural network for visual decision-making

Jie Su a,1, Fang Cai a,1, Shu-Kuo Zhao b, Xin-Yi Wang a, Tian-Yi Qian a,2, Da-Hui Wang b,2, Bo Hong a,2
PMCID: PMC12799151  PMID: 41512039

Significance

Biologically inspired neural networks offer interpretability but often underperform deep learning models due to limited optimization strategies. Here, we developed a model inspired by the primate dorsal visual pathway and introduced a neuroimaging-guided fine-tuning method. By leveraging correlations between human neuroimaging data and behavioral performance, this approach maps these findings onto the model’s key parameters and significantly enhances model performance while preserving biological plausibility, exhibiting human-like decision-making with robust resistance to noise and damage. Our work uniquely combines insights from primate electrophysiology and human neuroimaging, narrowing the gap between neuroimaging evidence and conventional artificial neural networks. It introduces a paradigm where neuroimaging directly guided the AI network design, enabling the creation of adaptive and resilient AI systems rooted in biological intelligence.

Keywords: neural dynamics model, spiking neural network, perceptual decision-making, MRI, neuroimaging-guided fine-tuning

Abstract

The human brain excels at complex tasks with remarkable efficiency, adaptability, and resilience, making it a powerful source of inspiration for AI. Here, we present a neural dynamics model inspired by the primate dorsal visual pathway, a circuit crucial for motion and spatial processing. Incorporating key neuronal and synaptic dynamics, the model reproduces human-like decision-making behaviors and neural activity patterns without the need for extensive training. Compared with conventional artificial networks, it exhibits superior robustness to perturbations such as noise and damage. To further enhance its performance, we introduce a neuroimaging-guided fine-tuning strategy. Correlations between MRI features and behavioral performance are mapped onto critical model parameters, guiding the optimization toward more biologically plausible operational regimes. This approach improves performance and adaptability while preserving biological plausibility and reducing the parameter search space. This is a demonstration of directly integrating human neuroimaging evidence into AI model optimization, establishing a methodology for brain-inspired modeling. By combining insights from primate electrophysiology, human neuroimaging, and biologically grounded modeling, our work narrows the gap between neuroscience and AI. It demonstrates how brain-inspired approaches can advance the development of adaptive, resilient, and interpretable AI systems, offering a paradigm for biologically grounded intelligence.


Abstracting the structures of biological neural systems into mathematical models to construct artificial neural networks for solving real-world problems is a pivotal approach to innovation in artificial intelligence. Over the past decades, this approach has achieved significant success in various fields (1–4). Visual perception, a fundamental process through which humans and animals interpret and interact with the environment, is a central topic in both neuroscience and AI. Understanding the neural mechanisms underlying perceptual decision-making not only provides insights into biological systems but also drives advancements in AI technologies. Convolutional neural network (CNN) models, inspired by the receptive fields and parallel distributed processing in animal vision systems, have achieved notable success in tasks such as object detection, facial recognition (5–7), and action recognition (8, 9), even surpassing human performance in some domains. However, these models face significant limitations, including the need for extensive labeled data, lack of biological interpretability, and susceptibility to adversarial attacks. These limitations highlight the gap between current AI systems and biological neural networks regarding robustness, efficiency, and flexibility. Addressing these challenges requires a paradigm shift toward models that more closely mimic the underlying principles of biological neural computation.

Motion perception is critical for animals to detect potential conspecifics, prey, or predators and is vital for survival. The motion perception system starts with retinal input, travels through the lateral geniculate nucleus (LGN) to the primary visual cortex (V1), and then projects to parietal areas along the dorsal visual pathway, supporting spatial attention and eye movements (10). Direction-selective (DS) neurons, which are essential for motion perception, are widely present in brain areas such as V1, the middle temporal region (MT), and the lateral intraparietal region (LIP) (11–13). The direction preference of V1 DS neurons depends on the spatial pattern of LGN neurons to which they are connected (14, 15). Motion perception relies on the integration of both spatial and temporal dimensions of information (16). Spatial integration is achieved by the larger receptive fields of MT neurons (17, 18), while temporal integration primarily relies on the temporal integration properties of LIP neurons (11, 19).

Based on the physiological structures and properties of neural circuits in LIP, Wang proposed a recurrent neural circuit model (20). In the model, two groups of excitatory DS neurons are capable of integrating lower-level synaptic inputs and performing cognitive tasks through recurrent excitation and mutual inhibition mechanisms. The model’s attractor dynamics enhance its ability to perform decision-making and working memory tasks by replicating neural activity patterns observed in nonhuman primates, which achieves performance consistent with experimental data. Subsequent studies extended the model through simplifications, theoretical analyses, and training methodologies (21–24). These studies primarily emphasize fitting and interpreting experimental data, showcasing biological plausibility and interpretability. However, model parameters are usually determined based on averaged experimental data, without considering the physiological characteristics that underlie behavioral differences among individuals. Moreover, these models have not been optimized using biological features derived from primate (including human) neuroimaging data, which limits their impact and application in AI.

Numerous studies have investigated the relationship between the physiological or structural features of the human brain and behavior (25). These studies span various human behaviors, including cognitive functions such as working memory (26), language acquisition (27), theory of mind (28). Modern neuroimaging technologies provide effective tools for quantifying human brain physiological features, including gray matter volume (29), cortical thickness (30), and myelination (31) calculated from structural images, measurements of fibers from diffusion images (32), and the correlation of blood-oxygen-level-dependent (BOLD) signals between brain regions in resting state, known as resting-state functional connectivity (rs-FC) (33–35). Alterations in white matter metrics have also been linked to neurodegenerative diseases (36–38). These studies have revealed correlations between brain structure, functional characteristics, and human behavior. Despite these insights, neuroimaging data have not yet been fully incorporated into the design or training of artificial neural networks (ANNs). Integrating such empirical findings with ANN modeling opens a promising pathway for developing brain-inspired AIs that more accurately replicate human behavioral performance.

Addressing the absence of biologically grounded tuning mechanisms that link behavioral and neuroimaging evidence to AI, this study synthesizes established neural principles to develop a comprehensive neural dynamics model of visual motion perception, offering avenues for biologically inspired AI. The model adheres to the physiological features of nonhuman primates and facilitates cognitive decision-making behavior in the random dot kinematogram (RDK) task, exhibiting biological-like behavioral and neuronal features. By comparing it with a CNN, we demonstrated that the model exhibits certain advantages in robustness against noise and damage. Furthermore, this study explores the correlation between behavioral and neuroimaging features of human participants during the RDK task through MRI techniques, particularly focusing on the structural and functional characteristics of corresponding brain regions in human experts. Finally, we introduce a neuroimaging-guided fine-tuning approach, which leverages these neuroimaging characteristics to optimize the biologically inspired neural network model and achieve notable improvements in performance.

Results

Neural Network Model that Mirrors Primate Behavior and Neural Activity.

We constructed a neural dynamics model encompassing the four key brain areas of the primate dorsal visual pathway (LGN, V1, MT, and LIP; Fig. 1A–C). This primate-informed neural network (PINN) is capable of performing motion perception and decision-making in the RDK task. The neurons in our model adopted the Leaky Integrate-and-Fire (LIF) model and were interconnected through excitatory synapses (AMPA, NMDA) and inhibitory synapses (GABA) (Materials and Methods: The Neural Dynamics Model). The spikes of different neuron groups were recorded during the RDK task (Fig. 1D–F).

Fig. 1.

Neural model for motion perception: V 1/M T/LIP circuit diagram and firing rate plots showing direction selectivity and coherence-based ramping.

A neural dynamics model of motion perception inspired by the dorsal visual pathway. (A) The RDK and the locations of the key structures within the dorsal visual pathway associated with motion perception and decision-making in the human brain. (B) Spatial characteristics and arrangements of LGN ON and OFF neurons, including their spatial and temporal profiles. (C) Structure of the biologically inspired neural dynamics model, illustrating four modules corresponding to those key brain regions and the connections between them. (D–F) Averaged firing rates of distinct groups of neurons in response to the RDK stimulus. Time 0 marks the stimulus onset. The solid line represents neuron groups with a preferred directional bias, while the dotted line represents neuron groups with an opposite bias. Color and legend indicate coherence levels. (D) V1 neurons demonstrate direction selectivity; (E) MT neurons enhance this selectivity, with the activation of direction-preference neurons increasing with coherence; (F) LIP neurons exhibit ramping activity and a winner-take-all effect, with the slope of ramping increasing as coherence levels rise.

Recordings showed that V1 neurons exhibit direction selectivity, which was further enhanced in the MT neurons. In the LIP module, neurons that prefer a specific motion direction gradually dominate, while those favoring the opposite direction are suppressed, resulting in a “winner-take-all” effect. Additionally, the population-averaged firing rates of both V1 and MT neurons increased with motion coherence and remained relatively stable throughout the stimulus presentation period. The rise in V1 population activity is likely attributed to an increase in the proportion of active neurons, whereas the increase in MT reflects elevated firing rates at the single-neuron level (SI Appendix, Fig. S2). In contrast, LIP neurons showed a gradual ramping of activation, with the ramping speed proportional to motion coherence (Fig. 1D–F). These simulated neural activities closely resemble those recorded in electrophysiological experiments with macaque monkeys (12, 39), demonstrating the biological plausibility of our model.

The performance of our model was evaluated using an RDK dataset (SI Appendix, Methods). The choice probability and average decision time (mapped from simulated steps) for each coherence level were calculated. We use the term “decision time” for the model, since nondecision components such as motor preparation and execution were not modeled. Fig. 2A (Upper panel) presents the psychometric curve of the model (in blue), which is similar to the psychometric curves of human participants (in orange). The model’s sensitivity (slope of the psychometric curve, k in SI Appendix, Eq. 1) is 19.31±0.17 (mean ± SEM of 45 repetitions), significantly exceeding the average level of human participants (36 participants, 15.1±1.9, t=2.50, P=0.015, two-sample t test). Moreover, the model’s decision time curve aligned with the reaction times of human participants. As coherence increases (task difficulty decreases), the decision time gradually becomes shorter, as illustrated in Fig. 2A (Lower panel). While the decision time can be extremely fast (around 100ms), it remains reasonable compared with human response times (less than 400 ms, SI Appendix, Fig. S3).

Fig. 2.

Figure showing model performance and sensitivity shifts (k) following electrical stimulation of M T, V 1, and LIP neurons during motion tasks.

Behavioral performance of the model in RDK tasks and effect of virtual electrical stimulation. (A) Behavioral performance of the model and human participants in the RDK task, showing selection probability, and average decision time or reaction time as functions of coherence level (upper: psychometric curve, lower: decision time curve). The model’s performance is represented in blue (average of 45 repetitions), and human performance is shown in orange (average of 36 participants). Error bars denote the SEM across all participants. The Inset subplot displays the distribution of sensitivity (k) for all participants (orange) and the model (blue, repeated 45 times). (B) Injecting current (0 to 30 pA) into left-preferring MT neurons shifts the psychometric and decision time curves rightward, decreasing sensitivity (regression coefficient β=−0.238, P<0.001) and shortening decision time. Color and colorbar represent the intensity of stimulation, consistent throughout. (C) Injecting current into right-preferring MT neurons shifts curves leftward, decreasing sensitivity (β=−0.266, P<0.001) and decision time. (D) Performance changes when additional electrical stimulation is applied to all neurons in the model’s V1. The additional input current decreases the slope of the psychometric curve (indicating reduced sensitivity, β=−0.605, P<0.001) and reduces mean decision time. (E) Sensitivity and decision time both decrease when additional electrical stimulation is applied to all neurons in MT (β=−0.265, P<0.001). (F) Applying the same intensity of current to LIP leads to similar effects. (β=−0.458, P<0.001).

In electrophysiological experiments, microstimulation of specific brain regions can induce behavioral changes (12, 39). To verify whether our PINN exhibits similar characteristics to biological motion perception systems, we conducted virtual cortical stimulation experiments. External currents were injected into different groups of neurons in our model to investigate if the results matched those observed in biological experiments (12, 39).

Applying continuous extra current (0 to 30 pA, Fig. 2B) to MT neurons that prefer leftward motion shifted the model’s psychometric curve to the right, indicating an increased leftward choice preference (β=−0.200, P<0.001). This manipulation slightly reduced the model’s sensitivity (β=−0.239, P<0.001), and decreased decision time for leftward motions but increased for rightward motions. Conversely, applying extra current to neurons that prefer rightward motion resulted in a leftward shift of the psychometric curve (intercept changed: β=0.194, P<0.001) and the decision time curve, while reducing sensitivity (β=−0.266, P<0.001; Fig. 2C). Stimu-lating the corresponding LIP neurons yielded similar results (SI Appendix, Fig. S4), aligning with observations in animal studies (39). Moreover, when we stimulated both groups of neurons in the MT region nonselectively (0 to 30 pA), the model’s sensitivity declined slightly (β=−0.265, P<0.001), but the decision time was notably reduced (Fig. 2E). Correspondingly, stimulating all neurons in the V1 region significantly reduced the model’s sensitivity (β=−0.605, P<0.001) and altered decision time (Fig. 2D). Stimulating all excitatory neurons in the LIP module led to effects similar to those reported in the literature (21), with both sensitivity and decision time decreasing (sensitivity: β=−0.458, P<0.001, Fig. 2F).

Model Optimization Inspired by Structural Connectivity of Human Experts.

Structural connectivity estimated from neuroimaging techniques reflects the mesoscopic properties of fibers, which are directly related to behavioral outcomes (25–28, 40). To identify the key parameters influencing the task performance, we collected behavioral and neuroimaging data from 36 human participants performing the RDK tasks.

In the correlation analysis, we found a significant negative correlation between the mean fractional anisotropy (FA) value of white matter in the left lateral occipital complex (LOC) region and participants’ perceptual threshold (a measure of noise tolerance at 79.4 % accuracy indicating behavioral performances; r=−0.435, P=0.024, Pearson correlation, FDR corrected, Fig. 3A). This indicates that participants with lower mean FA values in this region performed better in the RDK task. The left LOC, located in the occipital lobe, includes white matter pathways connecting V1 and MT. Lower FA values may reflect reduced anisotropy and predict a denser distribution of fiber bundles in this region.

Fig. 3.

Brain-model correlations: (A,C) task performance vs. F A and volume; (B,D) schematic and sensitivity plots for V 1-M T and M T-LIP connection tuning.

Structural feature correlation in the RDK task and neuroimaging-guided model tuning. (A) Mean FA values of the left LOC are negatively correlated with participants’ task performance (perceptual threshold). The red dot represents the best-performing participant (hereinafter the same). Participants who exhibit lower anisotropy (interpreted as higher fiber density) in this region performed better in the task. (B) Tuning the connection retention rate between V1 and MT in the model improves accuracy and efficiency. Upper: Schematic diagram of the adjusted V1–MT connection. Lower: Increasing the connection retention rate improves model sensitivity and reduces decision time. (C) The white matter volume in the right inferior parietal lobe, located between the occipital and parietal lobes, is positively correlated with task performance. This indicates that participants with better performance have larger white matter volume relative to total brain volume. (D) Tuning the proportion of connections between MT and LIP within a specific range improves the model’s accuracy and efficiency. Upper: Schematic diagram of the adjusted MT–LIP connection. Lower: Increasing the connection ratio initially enhances sensitivity and reduces decision time, but excessive connections can decrease sensitivity.

The relative white matter volume in the right inferior parietal lobe (IPL) region (i.e., the ratio of ROI volume to total brain volume, estimated from T1 images) also showed a marginal significant positive correlation with participants’ perceptual threshold (r=0.373, P=0.075, Pearson correlation, FDR corrected, Fig. 3C). This suggests that better-performing participants had larger white matter volume in this area, indicating more extensive fiber tracts. The right IPL region, located between the occipital and parietal lobes, includes the white matter pathway connecting the MT and LIP regions, which might be associated with decision-making in the brain. These results demonstrate correlations between perceptual threshold and the structural features of the dorsal visual pathway. No other structural features showed significant correlations.

Guided by these MRI findings, which implicated V1–MT and MT–LIP white matter connections in perceptual threshold, we employed our neuroimaging-guided fine-tuning approach to adjust corresponding model parameters. For the mean FA of left LOC region (Fig. 3A), we adjusted the connection retention rate of the V1–MT connections in the model (Fig. 3B, Upper panel) accordingly. Increasing this rate enhanced connection diversity, which we hypothesize to be associated with lower FA values. This modification enhanced model sensitivity and reduced decision time, aligning with the observed behavioral correlations (Fig. 3B). Similarly, adjusting the connection ratio between MT and LIP neurons (Fig. 3D, Upper panel) revealed that increasing the ratio from low levels improved model sensitivity and reduced decision time, aligning with neuroimaging and behavioral correlations. However, beyond a certain range, further increases in the connection ratio decreased both model sensitivity and decision time (Fig. 3D).

Model Optimization Inspired by Functional Connectivity of Human Experts.

The correlation between functional connectivity and perceptual threshold in human participants was also analyzed. We found a significant positive correlation between perceptual threshold and the resting-state functional connectivity (rs-FC) between the right MT and the right anterior agranular insular complex (AAIC, rh; r=0.445, P=0.040, Pearson correlation, FDR corrected, Fig. 4A). A stronger positive rs-FC between these regions was associated with better performance, potentially indicating enhanced self-monitoring.

Fig. 4.

Functional connectivity correlations and model tuning plots for M T and LIP synaptic conductance during motion perception.

Functional connectivity correlation in the RDK task and neuroimaging-guided model tuning. (A) Resting-state functional connectivity (rs-FC) between MT and AAIC in the right hemisphere is positively correlated with task performance (perceptual threshold). The positive connectivity may reflect enhanced self-monitoring during the task, leading to improved performance. (B) Increasing the average synaptic conductance in MT within a specific range improves the model’s sensitivity and efficiency. Upper: Increasing the synaptic conductance would essentially increase the V1–MT connection weights. Lower: As conductance increases, sensitivity initially rises and decision time decreases, but sensitivity eventually declines despite continued reduction in decision time. (C) Functional connectivity between LIP and the suborbital region in the right hemisphere is negatively correlated with task performance. The negative connectivity may reflect increased engagement during the task, which could lead to improved performance. (D) Increasing the average synaptic conductance in LIP within a specific range improves the model’s sensitivity and efficiency. Upper: Increasing the synaptic conductance would essentially increase the MT–LIP connection weights. Lower: Increasing the weight initially enhances sensitivity and reduces decision time, but excessive weight can decrease sensitivity. The Pearson correlation (FDR corrected) was used for these analyses; P-values are reported.

The rs-FC between the right LIP region and the right suborbital sulcus of the inferior frontal lobe (suborbital IFL, rh) was significantly negatively correlated with perceptual threshold (r=−0.472, P=0.032, Pearson correlation, FDR corrected, Fig. 4C, see also SI Appendix, Fig. S6). This suggests that a stronger negative correlation between resting-state activities in these regions is associated with better performance, possibly reflecting higher engagement or attentiveness during the task.

The functional MRI results highlighted brain areas not included in the existing dynamics model. These regions may provide top–down control to areas involved in motion perception and decision-making, thereby affecting perceptual decisions and behavior. Based on findings from electrophysiological experiments (41), we simulated the modulation of the target region by adjusting synaptic conductance in the corresponding module, without adding new brain regions to the model. The results are illustrated in Fig. 4 B and D. Increasing synaptic conductance in MT module enhanced model sensitivity and shortened decision time, mirroring the performance advantage linked to stronger MT–AAIC connectivity in humans (Fig. 4B). However, beyond a certain threshold, further increases in conductance led to a reduction in sensitivity, despite continued decreases in decision time. Similarly, elevating synaptic conductance in LIP resulted in a comparable pattern of effects (Fig. 4D).

These experiments illustrate how adjusting model parameters based on structural and functional features can enhance the neural dynamics model, reflecting biological behaviors observed in neuroimaging studies. This approach not only validates the model’s biological plausibility but also introduces a method for incorporating neuroimaging data into the fine-tuning of the neural dynamics model, which we term neuroimaging-guided fine-tuning.

Superior Robustness of Neural Dynamics Model Under Perturbation.

To assess the robustness of our model, we conducted four types of perturbation experiments on both the neural dynamics model (PINN) and a comparable CNN model. Perturbations were introduced by either dropping a fraction of connections or neurons, or by adding Gaussian noise to the connection weights or neural activations in each module of the models. Fig. 5B–I show the changes in accuracy of the CNN and PINN at each layer with varying perturbation intensities (see Table 1 for statistics; SI Appendix, Fig. S7 and Table S1 for changes in sensitivity).

Fig. 5.

Accuracy of C N N and P I N N models across four perturbations: connection dropout, weight noise, neuron dropout, and activation noise by model layer.

Perturbation comparison between CNN and the neural dynamics model. (A) The structure and parameters of the CNN model. (B–E) In the CNN model, changes in accuracy under different perturbation conditions: (B) when a certain percentage of connections are dropped in each layer; (C) when zero-mean Gaussian noise (with variance as a multiplier of the average absolute value) is added to the connection weights; (D) when a certain percentage of neurons are dropped; and (E) when zero-mean Gaussian noise is added to the neuron activations. Colors indicate the network layer subjected to perturbation (B and C, excluding the fully connected layer fc4) or the origin of connections (D and E). Cross marks indicate the accuracy of individual experiments, while the curves represent the overall accuracy averaged across five experiments. (F–I) In the PINN model, changes in accuracy under similar perturbation conditions: (F) when a certain percentage of connections are dropped in each layer; (G) when zero-mean Gaussian noise is added to the connection weights; (H) when a certain percentage of neurons are dropped; and (I) when zero-mean Gaussian noise is added to the neuron activations. Colors represent the neurons in each module (H and I) or the connections emanating from each module (F and G).

Table 1.

Regression on accuracy change vs. noise level

Noise type MotionNet β PINN β
Drop connection conv1 −0.530 LGN–V1 −0.458
conv2 −0.574 V1–MT −0.081
fc3 −0.361 MT–LIP −0.026
Add noise to connection conv1 −0.298 LGN–V1 −0.105
conv2 −0.263 V1–MT −0.270
fc3 −0.041 MT–LIP −0.006
Drop neurons conv1 −0.594 LGN −0.456
conv2 −0.411 V1 −0.179
fc3 −0.649 MT −0.089
fc4 −0.027 LIP −0.476
Add noise to neurons conv1 −0.004 LGN −0.304
conv2 −0.242 V1 −0.006
fc3 −0.122 MT −0.001
fc4 −0.083 LIP 0.004

Bold values reveal smaller changes in the performance drop.

As shown in these figures, our PINN model exhibited superior robustness compared to the CNN model under various perturbations. The CNN’s accuracy typically declined more significantly with increased perturbation. In contrast, the PINN model remained more stable, with its accuracy declining much more slowly in most cases (Table 1). Notably, the model’s performance remained largely unaffected when noise was introduced into modules corresponding to higher-level brain regions. For example, discarding connections from MT to LIP (Fig. 5F, red) or adding noise to these connections (Fig. 5G, red) had minimal impact on model accuracy, far less than the perturbation effects on the CNN (Fig. 5 B and C, red). Adding noise to the input current of neurons had almost no effect on PINN’s performance, except for perturbations to LGN neurons (Fig. 5I, the purple curve). Adding noise to LGN neurons caused model failure (Fig. 5 G and I, purple), likely due to the sparse connections between V1 and LGN leading to overreliance on LGN signals. Discarding neurons in the LIP module (Fig. 5H, blue) also drastically reduced model performance, as too few LIP neurons could not maintain the decision-related attractor states, causing the model to degrade to resting state with a single attractor (42), and thus unable to make decisions.

Additionally, we found that changes in accuracy and sensitivity with added noise in the PINN model were consistent, which contrasted with the behavior observed in CNN. In some cases, the CNN model retained high sensitivity even as its accuracy significantly decreased due to perturbation-induced bias. Furthermore, the CNN model exhibited greater variation with perturbation (see scatter distribution in Fig. 5B–D and SI Appendix, Figs. S8 and S9), while the PINN model showed smaller bias and maintained relatively stable performance even when disturbed. Unexpectedly, reducing the number of LIP neurons (SI Appendix, Fig. S7G, blue) and adding noise to LIP neuron input currents (SI Appendix, Fig. S7H, blue) in some cases improved model performance. This is because our base model was not fully optimized; similar perturbations that effectively weakened LIP recurrent connections or increased noise levels could thereby enhance performance.

Landscape of the Neural Dynamics Model.

To better understand the computational mechanism underlying decision-making, we constructed landscapes of the LIP module under various conditions using a statistical approach (42). The landscape “energy” U(r1,r2) was defined as the negative log probability of the steady state Pss, where lower values of U correspond to more probable states.

U(r1,r2)=−lnPss(r1,r2). [1]

In our PINN model, LIP neurons primarily rely on differences in input currents for decision-making. Larger input currents enable the model to reach the decision threshold more quickly but with reduced accuracy. For instance, increasing the connection strength between V1 and MT, between MT and LIP, or increasing the number of MT neurons connected to LIP enhances the total currents received by LIP neurons, driving the attractor of the model away from the resting state (Fig. 6A). This enhancement improves signal differentiation among neuronal groups that prefer different directions, thereby increasing accuracy (the initial rising phases in Fig. 3 B and D). However, larger input currents cause attractors in the decision space to shift toward the diagonal (Fig. 6B, from 2 to 3, see also example landscapes in SI Appendix, Fig. S10A). In this scenario, inhibitory neurons are less effective at suppressing DS neurons that prefer the opposite direction, disrupting the winner-take-all dynamics and deteriorating model performance (the subsequent declining phases in Fig. 3 B and D).

Fig. 6.

Impact of parameters on LIP decision space: (A) input current shifts, (B to C) energy landscape diagrams showing attractor states and transitions.

Diagram showing the impact of parameter adjustments on the model. (A) Adjusting connections between V1 and MT or MT and LIP alters the information received by LIP. Increasing connection weights shifts the input currents received by the two groups of direction-selective neurons in LIP from state 1 to state 2, thereby increasing both the differences in input currents between the two groups and their absolute magnitudes. (B) Changes in the information received by LIP affect the energy landscape and attractors in the decision space. Insufficient input currents fail to trigger a decision, maintaining LIP neurons at state 1; optimal input currents and current differences drive the state transition to decision state 2; excessively large input currents cause the state to shift toward 3, deviating from the intended decision state. (C) Changes in LIP’s recurrent connections also alter the landscape and attractor states in the decision space. Insufficient recurrent connections result in weak evidence being unable to drive LIP from resting state 1 to decision state 2; excessive recurrent connections cause the resting attractor state 1 to disappear and enlarge the attraction basins of decision states 2 and 3, where weak inputs or noise can drive LIP to one of these decision states.

Stronger recurrent connections eliminate the resting state attractor and expand the size of the decision state attractors (Fig. 6C, see also SI Appendix, Fig. S10B). Adjusting the number of LIP neurons leads to similar effects. If the parameters result in stable decision-state attractors, the model can make decisions and maintain the states without external stimuli, demonstrating a degree of working memory capability. This enables decision-making with limited information by integrating noisy evidence, which can shorten decision time but introduces a risk of errors. Conversely, if only a single resting state attractor exists, the model fails to reach a decision threshold and cannot make decisions, as evidenced by the performance collapse when LIP neurons were dropped (Fig. 5H, blue). In real-world scenarios, the model would not make a choice in this situation. Similarly, parameter adjustments can lead to another monostable mode, where neuron groups preferring opposite directions both exhibit high firing rates, resulting in an inability to make reasonable decisions (Fig. 5 G and I, purple).

Discussion

In this study, we developed a neural dynamics model (PINN) for motion perception and decision-making that simulates the dorsal visual pathway. The model incorporates neurons and synapses with dynamic characteristics and kinetic properties as its core elements. These components enable the model to exhibit neural activities and behavioral outputs that are comparable to those of the biological brain, including responses to virtual electrical stimulation. By leveraging neuroimaging insights from human experts with superior behavioral performance, we employed neuroimaging-guided fine-tuning to optimize the model parameters, leading to improved performance. Compared to CNNs, our model achieves similar performance with fewer parameters, more closely aligns with biological data, and exhibits greater robustness to perturbations.

Parameter optimization for neural dynamics models is both critical and challenging due to the complexity of nonlinear dynamical systems. Previous studies have primarily employed data fitting techniques using neural or behavioral data, which are feasible for small-scale models. However, as neural dynamics models increase in scale and complexity, with larger numbers of parameters and higher computational demands, optimization becomes significantly more computationally intensive and difficult. While it is theoretically possible to optimize key parameters in some simplified models through analytical methods, such methods are generally impractical for more complex models. This study explores the application of neuroimaging data in model optimization, a process we term neuroimaging-guided fine-tuning. While this approach may not achieve the global optimum, it offers practical directions for parameter optimization and significantly reduces the search space. It is noteworthy that excessive adjustments to parameters can impair model performance. This observation implies the existence of an optimal range for these parameters that maximizes performance. Biological neural systems, refined over billions of years of evolution, regulate these parameters within such an optimal range, with minor variations accounting for individual differences.

To implement neuroimaging-guided fine-tuning, we first identified brain regions associated with visual decision-making in human participants through behavioral and imaging experiments. Participants achieving superior performance in visual tasks tend to have lower mean FA values in the LOC region, indicating a higher degree of neural branching in the occipital lobe. Variations in white matter characteristics among the adults are thought to result from inherent brain structure and the long-term maturation process from childhood to adulthood (43), suggesting that individuals with advanced visual capabilities likely have undergone extensive training and development in visual functions. We observed asymmetric results in structural correlations. For instance, the mean FA value of the left LOC exhibited a significant correlation with perceptual threshold, whereas the correlation in the right LOC was not significant. A similar trend was observed in the right LOC but did not survive multiple comparisons correction (r=0.336, P=0.045, uncorrected; SI Appendix, Fig. S5B). Additionally, the relative white matter volume of the right, but not the left, IPL correlated with perceptual threshold (SI Appendix, Fig. S5D). We attribute this asymmetry to the influence of various developmental factors on structural features, which can lead to weak and lateralized correlations with performance on a specific task. We also identified correlations between perceptual threshold and the functional connectivity of MT–AAIC and LIP–suborbital IFL pairs. Significant asymmetric correlations emerged between the right AAIC and MT, and the right suborbital IFL and LIP (SI Appendix, Fig. S6). We interpreted this as evidence of top–down modulations from attention or working memory-related brain regions to the modeled areas—a finding consistent with the right-lateralization of such functions reported in previous studies (44, 45). Variations in functional connectivity between participants are primarily influenced by their level of participation and cognitive engagement during experiments (46), suggesting that expert participants are more engaged and attentive during tasks. We hypothesize that this process involves top–down regulation from higher-level brain areas to task-related regions, potentially enhancing synaptic transmission efficiency. The structural and functional connectivity features observed in human participants account for their superior performance from the perspectives of innate structure, long-term development, and short-term participation. These findings, applied to key parameter adjustments in the model, indicate that the neural dynamics model aligns with biological intelligence and shows good interpretability in physiological terms. This work proposes a potential path for enhancing model performance through neuroimaging-guided fine-tuning.

In 2-alternative forced choice (2-AFC) tasks, decision-making is often modeled as evidence accumulation over time, using approaches such as the drift diffusion model (DDM) (47–49), racing diffusion model (50, 51), and linear ballistic accumulator (LBA) model (52). The temporal integration in the LIP module resembles the DDM (see discussions in refs. 20–22 and 53), and can be directly related to it (54, 55). This type of evidence accumulation is not present in feedforward DCNN models. Additionally, the dynamic nature of the LIP constitutes an attractor network (21), where, through iterative processes, the network’s state stabilizes near attractors. This differs fundamentally from traditional DCNNs, which categorize through spatial partitioning within representational spaces, and may explain why neural dynamics models exhibit better stability and noise resistance.

As a proof of concept, the current model implements motion discrimination based on a specific arrangement of ON and OFF neurons to detect global motion in two directions (Section). Although functionally minimalistic, it is already capable of simulating a range of behavioral and physiological phenomena. The model could be further enhanced by incorporating known biological mechanisms—such as integration information of V1 complex cell in MT (56), encoding of multiple motion directions, and representation of motion speed—to support richer motion discrimination functions. Based on these ideas, more visual cognitive functions, such as anomalous motion detection, can be implemented in PINN. Nevertheless, through this simplified functional model, we have demonstrated the feasibility of constructing functional neural dynamics models by drawing entirely on biological neural architectures. We have also established a technical pathway whereby macroscopic neuroimaging data can guide the parameter optimization of microscopic models. As neuroimaging technologies advance, permitting increasingly detailed observation of brain-wide connectivity and dynamics, the proposed methods of conducting and tuning neural networks with certain functions will play an essential role in guiding the development of next-generation AI systems. By incorporating biologically plausible architectures and dynamics, such models may potentially achieve enhanced efficiency, robustness, and interpretability, which are critical for bridging insights between neural computation and AI.

Biological systems are widely recognized for their superior energy efficiency, which is considered a key advantage of brain-inspired models over traditional deep learning algorithms. To evaluate this, we compared the computational load between our PINN model and MotionNet on the RDK task (SI Appendix, Table S2). Since our PINN model does not require training, it exhibits a notable advantage in this regard. However, during the inference phase, although each iteration of the PINN requires less computation, the overall decision-making process typically involves hundreds to thousands of iterations—depending heavily on the numerical algorithm and step size—resulting in a significantly higher total computational cost compared to the single forward pass required by deep learning models. This analysis reveals that under conventional computing architectures, neural dynamics models face computational efficiency challenges, and their related modeling and applications are considerably constrained by hardware limitations. This underscores the urgent need for advances in neuromorphic computing and brain-inspired hardware technologies.

Conventionally, recurrent neural networks (RNNs) are often employed to model decision-making circuits. In our study, however, a hybrid model—constructed by replacing the fully connected layers in MotionNet with an RNN—proved exceptionally difficult to train while maintaining structural and dimensional parity with our PINN (SI Appendix, Fig. S11). Despite multiple attempts, the model rarely converged to a functional state on the RDK task, and its accuracy consistently underperformed MotionNet. We therefore focus our reporting on comparisons with the CNN-based model. The training challenges of plain RNNs are well established (57), motivating subsequent architectures like GRU and LSTM. That said, we speculatively hypothesize—though without substantive evidence at this stage—that incorporating biological constraints or bioinspired modifications into RNNs may alleviate these issues, though this remains to be validated.

In summary, integrating neural dynamics into AI models can bridge the gap between biological intelligence and artificial systems, leading to the development of next-generation AI that is both interpretable and resilient like biological organisms. By leveraging advances in neuroscience to construct models with neural dynamics, we are able to create AI systems that better mimic biological behavior. The neuroimaging-guided fine-tuning method enhances performance while preserving the advantages of neural dynamics. This approach ensures both high performance and explainability, aligning with biological plausibility and computational efficiency. This synthesis between advanced AI techniques and in-depth neuroscientific insights represents a promising direction for future research and development in both fields.

Materials and Methods

The Neural Dynamics Model.

The neural dynamics model (PINN) constructed in this study simulates four key areas of the dorsal pathway involved in motion perception, corresponding to the LGN, V1, MT, and LIP regions of the motion perception decision system from primates (Fig. 1A). We did not directly model retina, ganglion cells, etc., in detail, but simplified them into temporal and spatial convolutions, directly mapping visual input to the total input current of LGN neurons as the model’s input.

The LGN module of the model was constructed based on previous research (14, 15). The LGN was divided into two groups of ON and OFF neurons implemented using the Leaky Integrate-and-Fire (LIF) model, each group containing 10,000 neurons, all with double Gaussian spatial receptive fields (Eq. 2, Fig. 1B, Middle).

A(x,y)=απσα2exp−x2+y2σα2−βπσβ2exp−x2+y2σβ2, [2]

where α=1, β=1, σα=0.0894, σβ=0.1259 (14). These neurons covered a visual angle of 0.35 ° (x,y∈[−0.175,0.175] in Eq. 2), corresponding to a 9×9 pixel area. They were alternately overlapped and arranged regularly in space (Fig. 1B, Top), covering the entire 300×300 field of view, forming a two-dimensional plane corresponding to the image space. ON and OFF neurons have different temporal profiles, with ON neurons responding slower than OFF neurons (approximately 10 ms, Eq. 3, Fig. 1B, Bottom).

K(t)=αt6τ07exp−tτ0−βt6τ17exp−tτ1, [3]

where τ0=3.66, τ1=7.16, α=1, β=0.8 for ON neurons, and α=1, β=1 for OFF neurons (14). Images of 300×300 pixels were convolved spatially with the 9×9 spatial convolution kernels. The stimulus (2 s, 120 frames) was interpolated into 1,000 frames through nearest-neighbor interpolation to match the time step, and a temporal convolution with a sliding time window of 160 ms was performed. At each time point, two sets of 100×100 convolution results were obtained, which served as the stimulus-induced current for each neuron in the ON and OFF neuron groups respectively, and together with Ornstein–Uhlenbeck noise (58) form the input current of LGN neurons.

The other three brain regions were composed of LIF neuron groups, with each neuron group connected to the neurons in the previous layer with specific structures and probabilities. The V1 module had a total of 5,000 neurons, divided into two groups, namely G1 and G2. Each V1 neuron received input from one ON cell and one OFF cell arranged in a specific pattern through AMPA synapses with equivalent synaptic conductance of g¯=60 n S. In the LGN projection received by the G1 group neurons, the ON cells were always located to the left of the OFF cells. Similarly, in the LGN projection received by the G2 group neurons, the ON cells were always located to the right of the OFF cells (Fig. 1C). The MT module contained two groups of neurons, L and R, each consisting of 400 neurons. The L group neurons received input from the G1 group neurons within a certain receptive field, while the R group neurons received input from the G2 group neurons within the same receptive field. The connections between MT and V1 corresponding neuron groups were also formed by AMPA synapses, with a synaptic conductance of g¯=2.0 n S.

The LIP module was constructed according to the model in the literature (20), consisting of excitatory neuron groups A and B (each with 300 neurons) and an inhibitory neuron group I (500 neurons). The A group neurons received random projections from L, while the B group neurons received projections from R (50 % of neurons for each group) with AMPA synapses. Connections within excitatory groups A, B, as well as connections from excitatory groups to inhibitory neuron group I, were formed by AMPA and NMDA synapses with different temporal characteristics. The inhibitory neuron group I inhibited neurons in groups A and B through GABA synapses. The connection strengths (synaptic conductance coefficients) of the above connections followed a truncated normal distribution N^(g¯,0.5g¯,a=0), where the average conductance from MT to LIP was g¯MT=0.1 n S, the average conductance among excitatory neurons in LIP was g¯AMPA=0.05 n S, g¯NMDA=0.165 n S, the average conductance from excitatory neurons to inhibitory neurons in LIP was g¯AMPA=0.04 n S, g¯NMDA=0.13 n S. Additionally, the connection weights between neurons with the same direction preference increased to w+=1.3 times the original weight (Hebb-strengthened weight), while the connection weights between neurons with different direction preferences weakened to w−=0.7 times the original weight (Hebb-weakened weight).

The neurons in the model are all LIF neurons, a commonly used model in computational neuroscience to approximate the behavior of biological neurons. The resting membrane potential was set to Vr=−70 mV, and when the membrane potential of a neuron reaches −50 mV, it fires an action potential, resetting to −55 mV during the refractory period (59). The LIF model was selected due to its simplicity and computational efficiency, while still capturing essential dynamics of neuronal firing. In addition to synaptic currents and external currents, each neuron received Ornstein–Uhlenbeck noise, a standard approach for modeling stochastic fluctuations in neuronal input. The time constant of OU noise τ=10 ms, and mean current of 400 pA with a variance of 100 were chosen to match the noise characteristics observed in biological neurons. for LIP excitatory neurons, a higher mean current of 550 pA was used to simulate the enhanced excitatory drive from other excitatory neurons without direction selectivity. For excitatory neurons, the parameters were Cm=0.5 nF, gl=25 n S, and refractory period τ=2 ms. For inhibitory neurons, the parameters were Cm=0.2nF, gl=20 n S, and refractory period τ=1 ms. The AMPA, NMDA, and GABA synapses in the model followed the settings in the literature (20, 60).

We use the overall firing rate of neuron groups A or B in the LIP region as the basis for whether the model makes a decision. When the average firing rate exceeded a threshold (30 Hz) or when the stimulus finished, the model selected the direction preferred by the neuron group with the higher firing rate as its final output. By simulating the model with each stimulus in the aforementioned RDK dataset, we obtained the accuracy of the model’s decisions and the number of time steps taken for the decision (decision time) under different coherence levels. We then estimated a psychometric function through least squares regression. This method yields behavioral metrics, including the model’s psychometric curve and sensitivity indices.

Furthermore, our PINN model, which possesses a structure akin to dorsal visual pathway and neurons that mirror those of the biological nervous system, allows for the execution of virtual electrophysiological experiments, recording and analyzing the firing characteristics of each neuron group in the model, and even performing virtual electrical stimulation on neurons by adding additional current inputs, to study the characteristics of the model (Fig. 2 B and C).

Due to the randomness of RDK stimuli and the neural dynamics of the model, to ensure the robustness of the results, each stimulus was repeated twice in the model performance test, and the PINN model was reinitialized and repeated five times. The selection probabilities and average decision time for each coherence level were calculated and estimated using the psychometric curve, and the decision time curve was estimated using a moving median smoothing algorithm with a coherence window of 10 % (Fig. 2A, Upper panel). For human participants, we conducted a bootstrap analysis on the behavioral data, where the trials of each participant were randomly sampled with replacement, and the median of 1,000 bootstrap samples was calculated for each participant. The final average performance over all participants was plotted (Fig. 2A, orange).

A CNN Model for Motion Perception.

A CNN model (MotionNet) was built with structure and scale similar to the PINN model for comparison (Fig. 5A). The model received video input with a resolution of 300×300 pixels, processed it through spatial convolutions of 9×9 pixels and temporal convolutions of 10 frames, resulting in two 100×100 feature maps matching the input of LGN neurons. These feature maps were converted to two 50×50 maps using 3×3 convolutional kernels. The two channels were then processed with 11×11 convolutional kernels to generate two 20×20 feature maps. After average pooling along the time dimension, these features were mapped to a 400-dimensional vector by a linear layer, and finally to output neurons using another linear layer. All neurons utilized the ReLU activation function except the output layer, which was passed through a softmax transformation for fitting the one-hot encoded direction classification information.

To avoid overfitting the training data and ensure accurate model comparison, MotionNet was not directly trained on the RDK dataset. Instead, it utilized generated moving random images as the training data. Specifically, a random grayscale image was generated and stretched randomly. A window covering the image was moved in random directions to obtain the necessary animation data. The speed and direction of movement were randomly assigned, and the training labels (supervision information, moving left or right) were determined by the direction of movement along the horizontal axis.

MotionNet was trained with the cross-entropy loss function and stochastic gradient descent (SGD) method. The initial learning rate was set at 0.01, reduced to 10% at the 5th and 15th epochs. The momentum was set to 0.9, with a batch size of 64, and each epoch contained 500 batches. Training stopped at the 20th epoch. The model from the 20th epoch was selected for testing based on its convergence and stability during training. Similar analyses were then conducted on this model using the RDK dataset, following the same protocol as applied to the PINN model. This approach ensured a fair comparison between the two models’ performance under the same experimental conditions. Specifically, 20 consecutive frames were randomly selected from a 120-frame animation as input for MotionNet. The model’s choices under various coherences were recorded and psychometric curves were fitted. For result stability, each trial was repeated twice. Unlike the PINN model, MotionNet, being a CNN, lacks a concept of time, so we mainly focused on accuracy and sensitivity.

Neuroimaging Data Analysis Methods.

The study was approved by the Tsinghua University Science and Technology Ethics Committee (Medicine, THU01-20230217), and informed consent was obtained from each volunteer. The RDK behavioral and neuroimaging experiments are described in details in SI Appendix. High-resolution T1 structural images were parcellated into gray matter and white matter and reconstructed into cortex surfaces using FREESURFER (61). A voxel-level segmentation of brain regions based on the Destrieux atlas template (62) was also obtained. Several metrics were estimated, including the number of voxels per subregion, cortical area, and average cortical thickness. These processed data and structural features were used in the subsequent correlation analysis.

The fMRI data were preprocessed using FSL, including slice timing correction, motion correction, and field-map correction. The preprocessed images were then aligned to the individual’s T1 structural image. The WM/GM/CSF masks were constructed in the structural space for the subsequent GLM analysis to remove white matter and cerebrospinal fluid voxels. In the first-level analysis, the GLM model was used, and the rest, stimuli, and response phase of each trial were considered as three different regressors. Six motion parameters (three translation and three rotation parameters per frame) and the white matter and cerebrospinal fluid temporal signals were used as nuisance regressors in the model. The GLM model was fitted to each voxel, and the contrast of “stimuli − rest” was constructed to obtain the activation map in z-score. In the second-level analysis, each participant’s activation map in z-score was projected to the individual cortical space (reconstructed cortical surface), and then projected to a standard cortical space (fsaverage5 template), where the averaged activation map in z-score was obtained by performing a one-sample t test (P<0.001, uncorrected, SI Appendix, Fig. S1C).

The rs-FC in this study was obtained from the task-fMRI data (63). By regressing out the task-dependent signals and filtering the preprocessed results in the 0.01 to 0.08 Hz band, the background signals of the brain were obtained. The rs-FC matrix was then calculated using two atlases: the Destrieux structural atlas and the HCP extension atlas (64, 65). Matrix elements were used to characterize each ROI pair to correlate with the behavioral results.

Data preprocessing steps for diffusion images included brain extraction, registration to T1 images, field-map top-up correction, and eddy-current correction. We applied a ball-and-stick model (using FSL) to model the white matter voxels. The model regression of voxels in the white-matter region inferred the white-matter orientation distribution of single voxels, and FA and MD values for each voxel were then obtained (27). By registering to the white-matter segmentation image, white-matter regions’ mean FA and mean MD values (Destrieux atlas, threshold: FA > 0.2) were calculated for the subsequent correlation analysis. Mean FA and MD values were used to characterize the structural connectivity of white-matter pathways to correlate with the behavioral results.

To visualize the fiber bundles of the participants, fiber tract reconstruction was performed at the individual level using the quantitative anisotropy of the DSI Studio’s gqi inference as a tracking index. Segmented ROIs (LOC, IPL) were used as spatial constraints for fiber tracking. Deterministic fiber tracking was performed by the streamline Euler method (tracking threshold = 0, angular threshold = 0, min length = 30 mm, max length = 300 mm, seed number = 100,000) to obtain the morphology of fiber bundles related to the visual decision-making brain regions in the occipital and parietal lobes of the participants.

Neuroimaging-Behavioral Correlation Analysis Methods.

For correlation analysis between the structural MRI metrics and the behavioral results, we focused on brain ROIs required for modeling (i.e., LGN, V1, MT, LIP, including the gray matter regions and the white matter segments along the dorsal visual pathway) since the structural metrics could be related to the specific parameters used in the model (e.g., gray matter volumes vs. neuron number implemented in the model, mean FA values vs. connectivity probability in the model, etc.). The neuroimaging metrics we selected are commonly used in sMRI and dMRI data analyses, including: gray matter volume of each region, cortical surface area, mean cortical thickness, white matter volume (and its proportion to the whole-brain volume), mean fractional anisotropy (mean FA), mean diffusivity (mean MD). These parameters have been commonly used in studies of brain–behavior associations (25, 29–31).

For correlation analysis between the functional MRI metrics and the behavioral results, we focused on brain ROI pairs between modeled ROIs (LGN, V1, MT, LIP) and the ROIs with reported functional correlations. RDK task-related literature has reported correlations between behavioral performance and the electrophysiological properties in the inferior frontal gyrus, the insula cortex, and the posterior parietal cortex (66–68)—attributed to interparticipant variability in higher cognitive functions like working memory and visual attention. Our predefined functional ROIs included the inferior frontal lobe (IFL) and the insula cortex [i.e., subregions of these ROIs from the HCP extension atlas (65) and Destrieux atlas (62)].

The performance of the participants in the behavioral RDK experiment was quantitatively analyzed using psychometric curves, and Pearson correlation analysis was performed between the behavioral results and the structural and functional metrics estimated from the MRI data. For the behavioral data, a psychometric curve was fitted to each participant’s staircase performance. Behavioral performance was quantified using a noise tolerance metric as perceptual threshold, defined as the noise level (1 − coherence) at which participants maintained 79.4% accuracy. Higher values indicate superior performance on the RDK task. This metric is more sensitive to individual differences in perceptual ability than the raw percentage correct in our task.

Pearson correlations were calculated between the above metrics from structural (number of voxels per subregion, cortical area, average cortical thickness), diffusion (mean FA, mean MD), and functional (rs-FC) images and the perceptual threshold of the participants. The ROIs were determined by the parcellation and registration to the HCP extension atlas and the Destrieux atlas in each individual space. We calculated the correlations between the ROIs’ metrics and the behavioral perceptual threshold. We also examined the correlation with behavioral sensitivity, but no significant results were found (SI Appendix, Fig. S12). In contrast, Behavioral sensitivity showed a strong correlation with perceptual ability (r=−0.555, SI Appendix, Fig. S13). These differences may arise from the limited number of trials, which could lead to unreliable sensitivity estimates, or from participants adopting different strategies during the task.

Neuroimaging-Guided Fine-Tuning.

Connecting physiological metrics found in MRI with the model is challenging due to the different scales of neural elements involved. In this study, a heuristic approach was used to adjust and assess parameters potentially correlated with structural and functional MRI metrics, inspired by functional analogy. White matter connections found in MRI were simulated by adjusting the mean of the connection parameter distribution for V1 to MT (0.2 to 3.0% with a step size of 0.2), or by adjusting the proportion of connections between MT and LIP (10 to 100% with step 10%). According to Dale’s principle, nonpositive connection weights were set to zero, equivalent to removing those connections. The adjusted model was retested on the RDK dataset five times, and statistical analysis was performed on the sensitivity of psychometric curves (k) to investigate the impact of parameter adjustments on model performance and identify parameter combinations that improve model performance. Similarly, for functional connections found in fMRI, the synaptic conductance of neurons in MT was altered (0.01 to 0.15 with a step size of 0.01), or the Hebb-strengthened weight between LIP excitatory neurons was changed (0.7 to 1.5 with a step size of 0.1) to simulate the neural modulation from other brain regions. The same statistical analysis was conducted to examine the effect of parameter adjustments on model performance.

The relationship between model parameters (e.g., stimulation current, connection strength) and behavioral metrics (e.g., sensitivity) was quantified using linear regression. The regression coefficient (β) and the corresponding p-value were reported.

Perturbation Resistance Experiment of the Model.

In addition to requiring large amounts of data for training, deep learning models are also prone to biases and are easily affected by noise, resulting in poor transferability, among other issues. To address these problems, we designed a series of perturbation experiments on the neural dynamics model to assess its noise resistance performance.

For each module of the PINN model, we separately tested the effects of deactivating a certain percentage of neurons, discarding a certain percentage of synapses (ranging from 0 to 90% with a step size of 10%), adding a certain level of Gaussian perturbation to the input current of neurons (variance ranging from 0 to 2 times the averaged absolute value of the group’s input current, normal distribution noise with a mean of 0), and adding a certain level of Gaussian perturbation to the connection weights between neurons (variance ranging from 0 to 2 times the averaged absolute value of all connection weights, normal distribution noise with a mean of 0). For each parameter setting, the model’s behavioral performance was evaluated using the RDK dataset. Our analysis concentrated on understanding how changes in model parameters affected sensitivity and overall accuracy. This evaluation highlights the significance of each parameter and demonstrates the resiliency of our model.

For comparison, we conducted similar noise and damage experiments on MotionNet as we did on the PINN model. For each layer of MotionNet, we separately tested the effects of deactivating a certain percentage of neurons, discarding a certain percentage of connections (ranging from 0 to 90% with a step size of 10%), adding a certain level of Gaussian perturbation to the output of neurons (variance ranging from 0 to 2 times the averaged absolute value of the group’s neuron activation, normal distribution noise with a mean of 0), and adding a certain level of Gaussian perturbation to the connection weights between neurons (variance ranging from 0 to 2 times the averaged absolute value of all connection weights, normal distribution noise with a mean of 0). Due to the fact that the fully connected layer (fc4) of MotionNet is mapped to the output layer with only two neurons, the fc4 connections were neither deactivated nor perturbed during our experiment. After each adjustment of the perturbation parameters, we conducted a complete evaluation on the RDK dataset and calculated the sensitivity of the current model to the parameter changes.

Supplementary Material

Appendix 01 (PDF)

Movie S1.

Real-Time Visual Decision-Making by an Agent Based on the Neural Dynamics Model. The video showcases a visual decision-making agent utilizing our primate informed neural network model to perform motion perception task. The random dot kinematogram stimuli were presented on a screen, which the agent processed by capturing visual input through a high-speed camera and making decisions using two laser pointers. Throughout the task, the agent’s neural activities were displayed on a monitor, offering real-time insights into the dynamics of its decision-making processes.

Download video file (10MB, mp4)

Acknowledgments

This work was supportted by National Natural Science Foundation of China under grant 32171094. We gratefully acknowledge Hua-Shuo Liu, Jie-Ying Zhang, and Qing Li for their assistance with the human experiments.

Author contributions

J.S., T.-Y.Q., D.-H.W., and B.H. designed research; J.S., F.C., S.-K.Z., and X.-Y.W. performed research; J.S. and F.C. analyzed data; J.S., F.C., and X.-Y.W. prepared figures and SI Videos; and J.S., F.C., T.-Y.Q., D.-H.W., and B.H. wrote the paper.

Competing interests

The authors declare no competing interest.

Footnotes

This article is a PNAS Direct Submission. X.-J.W. is a guest editor invited by the Editorial Board.

Contributor Information

Tian-Yi Qian, Email: qiantianyi@qiyuanlab.com.

Da-Hui Wang, Email: wangdh@bnu.edu.cn.

Bo Hong, Email: hongbo@qiyuanlab.com.

Data, Materials, and Software Availability

Source code for the neural dynamics model (PINN), the convolutional neural network model (MotionNet), and the perturbation experiments are publicly available at GitHub (https://github.com/NeuralModeling/VisMotion). The RDK dataset, MRI data have been deposited in Zenodo (https://doi.org/10.5281/zenodo.17302015). All other data are included in the manuscript and/or supporting information.

Supporting Information

References

  • 1.McCulloch W. S., Pitts W., A logical calculus of the ideas immanent in nervous activity. Bull. Math. Biophys. 5, 115–133 (1943). [PubMed] [Google Scholar]
  • 2.Hopfield J. J., Neural networks and physical systems with emergent collective computational abilities. Proc. Natl. Acad. Sci. U.S.A. 79, 2554–2558 (1982). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3.LeCun Y., Bengio Y., Hinton G., Deep learning. Nature 521, 436–444 (2015). [DOI] [PubMed] [Google Scholar]
  • 4.Hassabis D., Kumaran D., Summerfield C., Botvinick M., Neuroscience-inspired artificial intelligence. Neuron 95, 245–258 (2017). [DOI] [PubMed] [Google Scholar]
  • 5.A. Krizhevsky, I. Sutskever, G. E. Hinton, Imagenet classification with deep convolutional neural networks. Commun. ACM 60, 84—90 (2017).
  • 6.C. Szegedy et al. , “Going deeper with convolutions” in 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (IEEE, Boston, MA, 2015), pp. 1–9.
  • 7.K. He, X. Zhang, S. Ren, J. Sun, “Deep residual learning for image recognition” in 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (IEEE, Las Vegas, NV, 2016), pp. 770–778.
  • 8.Ji S., Xu W., Yang M., Yu K., 3D convolutional neural networks for human action recognition. IEEE Trans. Pattern Anal. Mach. Intell. 35, 221–231 (2013). [DOI] [PubMed] [Google Scholar]
  • 9.D. Maturana, S. Scherer, “VoxNet: A 3D convolutional neural network for real-time object recognition” in 2015 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) (IEEE, Hamburg, Germany, 2015), pp. 922–928.
  • 10.Mishkin M., Ungerleider L. G., Contribution of striate inputs to the visuospatial functions of parieto-preoccipital cortex in monkeys. Behav. Brain Res. 6, 57–77 (1982). [DOI] [PubMed] [Google Scholar]
  • 11.Shadlen M. N., Newsome W. T., Neural basis of a perceptual decision in the parietal cortex (area lip) of the rhesus monkey. J. Neurophysiol. 86, 1916–1936 (2001). [DOI] [PubMed] [Google Scholar]
  • 12.Katz L. N., Yates J. L., Pillow J. W., Huk A. C., Dissociated functional significance of decision-related activity in the primate dorsal stream. Nature 535, 285–288 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.Yates J. L., Park I. M., Katz L. N., Pillow J. W., Huk A. C., Functional dissection of signal and noise in MT and LIP during decision-making. Nat. Neurosci. 20, 1285–1292 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.Chariker L., Shapley R., Hawken M., Young L. S., A theory of direction selectivity for macaque primary visual cortex. Proc. Natl. Acad. Sci. U.S.A. 118, e2105062118 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Chariker L., Shapley R., Hawken M., Young L. S., A computational model of direction selectivity in macaque v1 cortex based on dynamic differences between on and off pathways. J. Neurosci. 42, 3365–3380 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Burr D., Thompson P., Motion psychophysics: 1985–2010. Vis. Res. 51, 1431–1456 (2011). [DOI] [PubMed] [Google Scholar]
  • 17.Amano K., et al. , Human neural responses involved in spatial pooling of locally ambiguous motion signals. J. Neurophysiol. 107, 3493–3508 (2012). [DOI] [PubMed] [Google Scholar]
  • 18.Nishida S., Kawabe T., Sawayama M., Fukiage T., Motion perception: From detection to interpretation. Annu. Rev. Vis. Sci. 4, 501–523 (2018). [DOI] [PubMed] [Google Scholar]
  • 19.Shadlen M. N., Newsome W. T., Motion perception: Seeing and deciding. Proc. Natl. Acad. Sci. U.S.A. 93, 628–633 (1996). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.Wang X. J., Probabilistic decision making by slow reverberation in cortical circuits. Neuron 36, 955–968 (2002). [DOI] [PubMed] [Google Scholar]
  • 21.Wong K. F., Wang X. J., A recurrent network mechanism of time integration in perceptual decisions. J. Neurosci. 26, 1314–1328 (2006). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Wang X. J., Decision making in recurrent neuronal circuits. Neuron 60, 215–234 (2008). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23.Wang X. J., Macroscopic gradients of synaptic excitation and inhibition in the neocortex. Nat. Rev. Neurosci. 21, 169–178 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24.Song H. F., Yang G. R., Wang X. J., Reward-based training of recurrent neural networks for cognitive and value-based tasks. eLife 6, e21492 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25.Genon S., Eickhoff S. B., Kharabian S., Linking interindividual variability in brain structure to behaviour. Nat. Rev. Neurosci. 23, 307–318 (2022). [DOI] [PubMed] [Google Scholar]
  • 26.Colom R., Jung R. E., Haier R. J., General intelligence and memory span: Evidence for a common neuroanatomic framework. Cogn. Neuropsychol. 24, 867–878 (2007). [DOI] [PubMed] [Google Scholar]
  • 27.Hämäläinen S., Sairanen V., Leminen A., Lehtonen M., Bilingualism modulates the white matter structure of language-related pathways. NeuroImage 152, 249–257 (2017). [DOI] [PubMed] [Google Scholar]
  • 28.Rice K., Redcay E., Spontaneous mentalizing captures variability in the cortical thickness of social brain regions. Soc. Cogn. Affect. Neurosci. 10, 327–334 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29.Mechelli A., Price C. J., Friston K. J., Ashburner J., Voxel-based morphometry of the human brain: Methods and applications. Curr. Med. Imaging 1, 105–113 (2005). [Google Scholar]
  • 30.Fischl B., Dale A. M., Measuring the thickness of the human cerebral cortex from magnetic resonance images. Proc. Natl. Acad. Sci. U.S.A. 97, 11050–11055 (2000). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31.Glasser M. F., Van Essen D. C., Mapping human cortical areas in vivo based on myelin content as revealed by T1-and T2-weighted MRI. J. Neurosci. 31, 11597–11616 (2011). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32.Pievani M., et al. , Assessment of white matter tract damage in mild cognitive impairment and Alzheimer’s disease. Hum. Brain Mapping 31, 1862–1875 (2010). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33.Sui J., Huster R., Yu Q., Segall J. M., Calhoun V. D., Function-structure associations of the brain: Evidence from multimodal connectivity and covariance studies. Neuroimage 102, 11–23 (2014). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34.Wang K., et al. , Altered functional connectivity in early Alzheimer’s disease: A resting-state fMRI study. Hum. Brain Mapping 28, 967–978 (2007). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 35.Song M., et al. , Brain spontaneous functional connectivity and intelligence. Neuroimage 41, 1168–1176 (2008). [DOI] [PubMed] [Google Scholar]
  • 36.Schmithorst V. J., Yuan W., White matter development during adolescence as shown by diffusion MRI. Brain Cogn. 72, 16–25 (2010). [DOI] [PubMed] [Google Scholar]
  • 37.Kantarci K., et al. , White-matter integrity on DTI and the pathologic staging of Alzheimer’s disease. Neurobiol. Aging 56, 172–179 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 38.Forkel S. J., Friedrich P., Thiebaut M., de Schotten H., Howells, White matter variability, cognition, and disorders: A systematic review. Brain Struct. Funct. 227, 529–544 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 39.Hanks T. D., Ditterich J., Shadlen M. N., Microstimulation of macaque area lip affects decision-making in a motion discrimination task. Nat. Neurosci. 9, 682–689 (2006). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 40.Kanai R., Bahrami B., Roylance R., Rees G., Online social network size is reflected in human brain structure. Proc. R. Soc. B: Biol. Sci. 279, 1327–1334 (2012). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 41.Briggs F., Mangun G. R., Usrey W. M., Attention enhances synaptic efficacy and the signal-to-noise ratio in neural circuits. Nature 499, 476–480 (2013). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 42.Ye L., Li C., Quantifying the landscape of decision making from spiking neural networks. Front. Comput. Neurosci. 15, 740601 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 43.Lebel C., Walker L., Leemans A., Phillips L., Beaulieu C., Microstructural maturation of the human brain from childhood to adulthood. NeuroImage 40, 1044–1055 (2008). [DOI] [PubMed] [Google Scholar]
  • 44.Bartolomeo P., Seidel Malkinson T., Hemispheric lateralization of attention processes in the human brain. Curr. Opin. Psychol. 29, 90–96 (2019). [DOI] [PubMed] [Google Scholar]
  • 45.Petit L., et al. , Strong rightward lateralization of the dorsal attentional network in left-handers with right sighting-eye: An evolutionary advantage. Hum. Brain Mapp. 36, 1151–1164 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 46.Küün S., Forlim C. G., Lender A., Wirtz J., Gallinat J., Brain functional connectivity differs when viewing pictures from natural and built environments using fMRI resting state analysis. Sci. Rep. 11, 4110 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 47.Ratcliff R., A theory of memory retrieval. Psychol. Rev. 85, 59 (1978). [DOI] [PubMed] [Google Scholar]
  • 48.Ratcliff R., Modeling response signal and response time data. Cogn. Psychol. 53, 195–237 (2006). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 49.Ratcliff R., McKoon G., The diffusion decision model: Theory and data for two-choice decision tasks. Neural Comput. 20, 873–922 (2008). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 50.Usher M., McClelland J. L., The time course of perceptual choice: The leaky, competing accumulator model. Psychol. Rev. 108, 550 (2001). [DOI] [PubMed] [Google Scholar]
  • 51.Tillman G., Van Zandt T., Logan G. D., Sequential sampling models without random between-trial variability: The racing diffusion model of speeded decision making. Psychon. Bull. Rev. 27, 911–936 (2020). [DOI] [PubMed] [Google Scholar]
  • 52.Brown S. D., Heathcote A., The simplest complete model of choice response time: Linear ballistic accumulation. Cogn. Psychol. 57, 153–178 (2008). [DOI] [PubMed] [Google Scholar]
  • 53.Wei H., Bu Y., Dai D., A decision-making model based on a spiking neural circuit and synaptic plasticity. Cogn. Neurodyn. 11, 415–431 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 54.Roxin A., Ledberg A., Neurobiological models of two-choice decision making can be reduced to a one-dimensional nonlinear diffusion equation. PLoS Comput. Biol. 4, e1000046 (2008). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 55.Umakantha A., Purcell B. A., Palmeri T. J., Relating a spiking neural network model and the diffusion model of decision-making. Comput. Brain Behav. 5, 279–301 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 56.Simoncelli E. P., Heeger D. J., A model of neuronal responses in visual area MT. Vis. Res. 38, 743–761 (1998). [DOI] [PubMed] [Google Scholar]
  • 57.R. Pascanu, T. Mikolov, Y. Bengio, “On the difficulty of training recurrent neural networks” in Proceedings of the 30th International Conference on Machine Learning/Proceedings of Machine Learning Research, S. Dasgupta, D. McAllester, Eds. (PMLR, Atlanta, GA, 2013), vol. 28, pp. 1310–1318.
  • 58.Lánskỳ P., Sacerdote L., The Ornstein–Uhlenbeck neuronal model with signal-dependent noise. Phys. Lett. A 285, 132–140 (2001). [DOI] [PubMed] [Google Scholar]
  • 59.Troyer T. W., Miller K. D., Physiological gain leads to high ISI variability in a simple model of a cortical regular spiking cell. Neural Comput. 9, 971–983 (1997). [DOI] [PubMed] [Google Scholar]
  • 60.Jahr C. E., Stevens C. F., Voltage dependence of NMDA-activated macroscopic conductances predicted by single-channel kinetics. J. Neurosci. 10, 3178–3182 (1990). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 61.Fischl B., Freesurfer. Neuroimage 62, 774–781 (2012). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 62.Destrieux C., Fischl B., Dale A., Halgren E., Automatic parcellation of human cortical gyri and sulci using standard anatomical nomenclature. Neuroimage 53, 1–15 (2010). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 63.Fox M. D., et al. , Combining task-evoked and spontaneous activity to improve pre-operative brain mapping with fMRI. NeuroImage 124, 714–723 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 64.Glasser M. F., et al. , A multi-modal parcellation of human cerebral cortex. Nature 536, 171–178 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 65.Huang C. C., Rolls E. T., Feng J., Lin C. P., An extended Human Connectome Project multimodal parcellation atlas of the human cortex and subcortical areas. Brain Struct. Funct. 227, 763–778 (2022). [DOI] [PubMed] [Google Scholar]
  • 66.Tops M., Boksem M. A. S., A potential role of the inferior frontal gyrus and anterior insula in cognitive control, brain rhythms, and event-related potentials. Front. Psychol. 2, 330 (2011). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 67.von Lautz A., Herding J., Blankenburg F., Neuronal signatures of a random-dot motion comparison task. NeuroImage 193, 57–66 (2019). [DOI] [PubMed] [Google Scholar]
  • 68.Uddin L. Q., Salience processing and insular cortical function and dysfunction. Nat. Rev. Neurosci. 16, 55–61 (2015). [DOI] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Appendix 01 (PDF)

Movie S1.

Real-Time Visual Decision-Making by an Agent Based on the Neural Dynamics Model. The video showcases a visual decision-making agent utilizing our primate informed neural network model to perform motion perception task. The random dot kinematogram stimuli were presented on a screen, which the agent processed by capturing visual input through a high-speed camera and making decisions using two laser pointers. Throughout the task, the agent’s neural activities were displayed on a monitor, offering real-time insights into the dynamics of its decision-making processes.

Download video file (10MB, mp4)

Data Availability Statement

Source code for the neural dynamics model (PINN), the convolutional neural network model (MotionNet), and the perturbation experiments are publicly available at GitHub (https://github.com/NeuralModeling/VisMotion). The RDK dataset, MRI data have been deposited in Zenodo (https://doi.org/10.5281/zenodo.17302015). All other data are included in the manuscript and/or supporting information.


Articles from Proceedings of the National Academy of Sciences of the United States of America are provided here courtesy of National Academy of Sciences

RESOURCES