Skip to main content
Proceedings of the National Academy of Sciences of the United States of America logoLink to Proceedings of the National Academy of Sciences of the United States of America
. 2023 Aug 28;120(36):e2221704120. doi: 10.1073/pnas.2221704120

A law of data separation in deep learning

Hangfeng He a,b, Weijie J Su c,1
PMCID: PMC10483613  PMID: 37639604

Significance

The practice of deep learning has long been shrouded in mystery, leading many to believe that the inner workings of these black-box models are chaotic during training. In this paper, we challenge this belief by presenting a simple and approximate law that deep neural networks follow when processing data in the intermediate layers. This empirical law is observed in a class of modern network architectures for vision tasks, and its emergence is shown to bring important benefits for the trained models. The significance of this law is that it allows for a perspective that provides useful insights into the practice of deep learning.

Keywords: deep learning, data separation, constant geometric rate, intermediate layers

Abstract

While deep learning has enabled significant advances in many areas of science, its black-box nature hinders architecture design for future artificial intelligence applications and interpretation for high-stakes decision-makings. We addressed this issue by studying the fundamental question of how deep neural networks process data in the intermediate layers. Our finding is a simple and quantitative law that governs how deep neural networks separate data according to class membership throughout all layers for classification. This law shows that each layer improves data separation at a constant geometric rate, and its emergence is observed in a collection of network architectures and datasets during training. This law offers practical guidelines for designing architectures, improving model robustness and out-of-sample performance, as well as interpreting the predictions.


Deep learning methodologies have achieved remarkable success across a wide range of data-intensive tasks in image recognition, biological research, and scientific computing (1–4). In contrast to other machine learning techniques (5), however, the practice of deep learning relies heavily on a plethora of heuristics and tricks that are not well justified. This situation often makes deep learning-based approaches ungrounded for some applications or necessitates the need for huge computational resources for exhaustive search, making it difficult to fully realize the potential of this set of methodologies (6).

This unfortunate situation is in part owing to a lack of understanding of how the prediction depends on the intermediate layers of deep neural networks (7–9). In particular, little is known about how the data of different classes (e.g., images of cats and dogs) in classification problems are gradually separated from the bottom layers to the top layers in modern architectures such as AlexNet (1) and residual neural networks (10). Any knowledge about data separation, especially quantitative characterization, would offer useful principles and insights for designing network architectures, training processes, and model interpretation.

The main finding of this paper is a quantitative delineation of the data separation process throughout all layers of deep neural networks. As an illustration, Fig. 1 plots a certain value that measures how well the data are separated according to their class membership at each layer for feedforward neural networks trained on the Fashion-MNIST dataset (12). This value in the logarithmic scale decays, in a distinct manner, linearly in the number of layers the data have passed through. The Pearson correlation coefficients between the logarithm of this value and the layer index range from −0.997 to −1 in Fig. 1.

Fig. 1.

Fig. 1.

Illustration of the law of equi-separation in feedforward neural networks with ReLU activations trained on the Fashion-MNIST dataset. The three rows correspond to three different training methods, stochastic gradient descent (SGD), SGD with momentum, and Adam (11). Throughout the paper, the x axis represents the layer index, and the y axis represents the separation fuzziness defined in Eq. 1, unless otherwise specified. The Pearson correlation coefficients, by row first, are −1.000, −0.998, −0.997, −0.999, −0.998, −0.998, −1.000, −0.997, and −0.999. The last two rows show the intermediate data points on the plane of the first two principal components for the 8-layer network trained using Adam. Layer 0 corresponds to the raw input. More details can be found in SI Appendix.

The measure shown in Fig. 1 is canonical for measuring data separation in classification problems. Let xki denote an intermediate output of a neural network on the ith point of Class k for 1≤i≤nk, x¯k denote the sample mean of Class k, and x¯ denote the mean of all n:=n1+⋯+nK data points. We define the between-class sum of squares and the within-class sum of squares as

SSb:=1n∑k=1Knk(x¯k−x¯)(x¯k−x¯)⊤SSw:=1n∑k=1K∑i=1nk(xki−x¯k)(xki−x¯k)⊤,

respectively. The former matrix represents the between-class “signal” for classification, whereas the latter denotes the within-class variability. Writing SSb+ for the Moore–Penrose inverse of SSb (The matrix SSb has rank at most K−1 and is not invertible in general since the dimension of the data is typically larger than the number of classes.), the ratio matrix SSwSSb+ can be thought of as the inverse signal-to-noise ratio. We use its trace

D:=Tr(SSwSSb+), [1]

to measure how well the data are separated (13, 14). This value, which is referred to as the separation fuzziness, is large when the data points are not concentrated to their class means or, equivalently, are not well separated, and vice versa.

1. Main Results

Given an L-layer feedforward neural network, let Dl denote the separation fuzziness (Eq. 1) of the training data passing through the first l layers for 0≤l≤L−1.* Fig. 1 suggests that the dynamics of Dl follows the relation

Dl≐ρlD0, [2]

for some decay ratio 0<ρ<1. Alternatively, this law implies logDl+1−logDl≐−log1ρ, showing that the neural network makes equal progress in reducing logD over each layer on the training data. Hence, we call this the law of equi-separation. This law is a quantitative and geometric characterization of the data separation process in the intermediate layers. Indeed, it is unexpected because the intermediate output of the neural network does not exhibit any quantitative patterns, as shown by the last two rows of Fig. 1.

The decay ratio ρ depends on the depth of the neural network, dataset, training time, and network architecture and is also affected, to a lesser extent, by optimization methods and many other hyperparameters. For the 20-layer network trained using Adam (11) (the bottom-right plot of Fig. 1), the decay ratio ρ is 0.818. Thus, the half-life is ln2lnρ−1=0.693lnρ−1=3.45, suggesting that this 20-layer neural network reduces the value of the separation fuzziness in every three and a half layers.

This law manifests in the terminal phase of training (14), where we continue to train the model to interpolate in-sample data. At initialization, the separation fuzziness may even increase from the bottom to top layers. During the early stages of training, the bottom layers tend to learn faster at reducing the separation fuzziness compared to the top layers (SI Appendix, Fig. S7). However, as training progresses, the top layers eventually catch up as the bottom layers have learned the necessary features. Finally, each layer becomes roughly equally capable of reducing the separation fuzziness multiplicatively. This dynamics of data separation during training is illustrated in Fig. 2. Neural networks in the terminal phase of training also exhibit certain symmetric geometries in the last layer such as neural collapse (14, 15) and minority collapse (9). However, it is worthwhile mentioning that the law of equi-separation emerges earlier than neural collapse during training (SI Appendix, Fig. S1).

Fig. 2.

Fig. 2.

A 20-layer feedforward neural network trained on Fashion-MNIST. The law of equi-separation starts to emerge at epoch 100 and becomes more clear as training proceeds. As in Fig. 1 and all the other figures in the main text, the x axis represents the layer index, and the y-axis represents the separation fuzziness.

The pervasive law of equi-separation consistently prevails across diverse datasets, learning rates, and class imbalances, as illustrated in Fig. 3. Additionally, SI Appendix, Fig. S6 demonstrates its applicability in a finer, class-wise context. This law is further exemplified in contemporary network architectures for vision tasks, such as AlexNet and VGGNet (16), as shown in Fig. 4 (see SI Appendix, Figs. S2 and S3 for additional convolutional neural network experiments). Moreover, the law manifests in residual neural networks and densely connected convolutional networks (17) when separation fuzziness is assessed at each block, as depicted in Figs. 5 and 6, respectively. Intriguingly, this law appears slightly more pronounced in feedforward neural networks when compared with other network architectures.

Fig. 3.

Fig. 3.

Illustration showing how the law of equi-separation holds in a wide range of settings. The experimental details are provided in SI Appendix.

Fig. 4.

Fig. 4.

Illustration of the equi-separation law on convolutional neural networks. A more in-depth examination of convolutional neural networks can be found in SI Appendix, Figs. S2 and S3.

Fig. 5.

Fig. 5.

Illustration of the law of equi-separation in residual neural networks: Each block contains either two or three layers. In the “Mixed” configuration, the first two blocks consist of three layers each, while the last two blocks have two layers each. In each plot, a layer or a block is identified as a module when presenting the x-axis. The first two columns are evaluated on the Fashion-MNIST dataset, and the last two columns are evaluated on the CIFAR-10 dataset.

Fig. 6.

Fig. 6.

Illustration of the law of equi-separation in densely connected convolutional networks (DenseNet161) by identifying a block as a module.

The separation dynamics of neural networks have been extensively investigated in prior research studies (18–22). For instance, ref. 18 employed linear classifiers as probes to assess the separability of intermediate outputs. In ref. 19, the author scrutinized the separation capabilities of neural networks through empirical spectral analysis. More recently, refs. 20–22 explored the separation ability of neural networks by examining neural collapse at intermediate layers and its relationship with generalization. In particular, ref. 22 provided crucial experimental evidence illustrating the progression of neural collapse within the interior of neural networks, which is perhaps the most relevant work to our paper.

2. Insights from the Law

Next, we demonstrate the benefits of the law of equi-separation and the perspective of data separation in regard to the three fundamental aspects of deep learning practice, namely, architecture design, training, and interpretation.

A. Network Architecture.

The law of equi-separation can inform the design principles for network architecture. First, the law implies that neural networks should be deep for good performance, thereby reflecting the hierarchical nature of deep learning (2, 23–25). This is because, as shown by Eq. 2, all layers reduce the separation fuzziness D0 of the raw input to ρL−1D0 at the last layer. When L is small—say, L=2 or 3—the ratio ρL−1 would generally not be small, and therefore, the neural network is unlikely to separate the data well. In the literature, the fundamental role of depth is recognized by analyzing the loss functions (10, 16, 23, 26, 27), and our law of equi-separation offers a perspective on network depth.

For completeness, the deeper the better is not necessarily true because as the depth L increase, ρ might become larger. Moreover, a very large depth would render the optimization problem challenging. As shown in Fig. 1, the 20-layer neural networks have higher values of the final separation fuzziness than those of their 8-layer counterparts. The Fashion-MNIST dataset is not very complex and, therefore, an 8-layer network will suffice to classify the data well. Adding additional layers, therefore, would only give rise to optimization burdens. Indeed, Fig. 7 illustrates the law of equi-separation with varying network depths, showing that different datasets correspond to different optimal depths. Therefore, as a practical guideline, the choice of depth should consider the complexity of the applications. For instance, this perspective of data separation shows that the optimal depths for MNIST (28), Fashion-MNIST, and CIFAR-10 (29) are 6, 10, and 12, respectively. This is consistent with the increasing complexity from MNIST to Fashion-MNIST to CIFAR-10 (10, 30, 31).

Fig. 7.

Fig. 7.

Illustration of the law of equi-separation with varying network depths. The legend shows how depths are color coded. The depth that yields the lowest separation fuzziness depends on the dataset complexity.

Moreover, the law of equi-separation indicates that depth bears a more significant connection to training performance than the “shape” of neural networks. In the first row of Fig. 8, the network does not separate the data well when the width of each layer is merely 20. This is because the network has to be wide enough to pass the useful features of the data for classification. However, when the network is wide enough, the network performance may be saturated even if more neurons are added to each layer (32). In addition, the second row of Fig. 8 shows that it is better to start with relatively wide layers at the bottom layers followed by narrow top layers. These observations suggest that very wide neural networks in general should not be recommended as they consume more time for optimization. Apart from encoding purposes, the width should be set to about the same across all layers or be made larger for the bottom layers (30, 33).

Fig. 8.

Fig. 8.

Illustration of the law of equi-separation with different widths and shapes of the neural networks. For example, “Narrow-Wide” means that the bottom layers are narrower than the top layers.

B. Training.

The emergence of the law of equi-separation during training is indicative of good model performance. For example, its manifestation can improve the robustness to model shifts. Indeed, the data separation ability of the network, as measured by DL−1D0, becomes least sensitive to perturbations in network weights when the law manifests. To see this point, let

R:=DL−1DL−2×DL−2DL−3×⋯×D1D0=DL−1D0,

denote the reduction ratio of the separation fuzziness made by the entire network. Assuming that a perturbation of the network weights leads to a change of ε in the ratio for each layer, the perturbed reduction ratio becomes

Rε:=DL−1DL−2+εDL−2DL−3+ε⋯D1D0+ε=R+RDL−2DL−1+DL−3DL−2+⋯+D0D1ε+O(ε2). [3]

Applying the inequality of arithmetic and geometric means, the perturbation term RDL−2DL−1+DL−3DL−2+⋯+D0D1ε is minimized in absolute value when DL−1DL−2=DL−2DL−3=⋯=D1D0. This result is summarized in the following proposition:

Proposition 1

For small ε, |Rε−R| is minimized when

DL−1DL−2=DL−2DL−3=⋯=D1D0.

This condition is precisely the case when the equi-separation law holds. Thus, the data separation ability of the classifier is not very sensitive to perturbations in network weights when the law emerges. Therefore, to improve the robustness, we need to train a neural network at least until the law of equi-separation comes into effect. In the literature, however, connections with robustness are often established through loss functions rather than the data separation perspective (34).

This proposition suggests that the law of equi-separation could be considered a demonstration of the neural networks’ “homogeneity” across all layers. An essential component in the derivation of Proposition 1involves the assumption that the magnitude of the perturbation remains uniform across all layers, as seen from Eq. 3. Even though it is improbable for the perturbation to remain constant in practice, it is intriguing to note that the emergence of this law is contingent on the use of batch normalization (35), which enhances the homogeneity of the networks in a certain sense. Indeed, SI Appendix, Fig. S10 shows that the law does not appear when batch normalization is not used.

The law of equi-separation can also shed light on the out-of-sample performance of neural networks. Fig. 9 considers two neural networks where one exhibits the law of equi-separation and the other does not. The results of our experiments showed that the former one has better test performance. Although this law is sometimes not observed in neural networks with good test performance, interestingly, one can fine-tune the parameters to reactivate the law with about the same or better test performance.

Fig. 9.

Fig. 9.

Illustration showing the law of equi-separation manifesting in the left network but not in the right network because of frozen training. Both networks have about the same training losses as well as training accuracies. However, the left network has a higher test accuracy (23.85%) than that of the right network (19.67%).

Moreover, the law of equi-separation can be potentially used to facilitate the training of transfer learning. In transfer learning, for example, bottom layers trained from an upstream task can form part of a network that will be trained on a downstream task. A challenge arising in practice is how to balance between the training times on the source and the target. In light of the law of equi-separation, a sensible balance is to find when the improvement in reducing separation fuzziness is roughly the same between the bottom and top layers.

C. Interpretation.

The equi-separation law offers a perspective for interpreting deep learning predictions, especially in high-stakes decision-makings. An important ingredient in interpretation is to pinpoint the basic operational modules in neural networks and subsequently to reflect that “all modules are created equal.” For feedforward and convolutional neural networks, each layer is a module as it reduces the separation fuzziness by an equal multiplicative factor (for more details, see SI Appendix, Fig. S11). Interestingly, even when the law no longer manifests in residual neural networks, Fig. 5 shows that it can be restored by identifying a block rather than a layer as a module. This figure also suggests that blocks containing more layers are more capable of reducing the separation fuzziness. Likewise, Fig. 6 shows that the law of equi-separation holds in densely connected convolutional networks by identifying a block as a module. For comparison, when the perspective of data separation is not considered (36), the correct module cannot be identified. For completeness, it should be noted that the law becomes less clear for deeper residual neural networks,† as shown in Fig. 10.

Fig. 10.

Fig. 10.

The law of equi-separation in ResNet18 is not as pronounced as it is in shallower residual neural networks.

Because each module makes equal but small contributions, a few neurons in a single layer are unlikely to explain how a deep learning model makes a particular prediction. It is therefore crucial to take all layers collectively for interpretation (37). This view, however, challenges the widely used layer-wise approaches to deep learning interpretation (38, 39).

3. Discussion

Due to ref. 14 and follow-up work, it has been known that neural networks can exhibit precise mathematical structures at the last layer in the terminal phase of training. In this paper, we have extended this viewpoint from the surface of these black-box models to their interior by introducing an empirical law that quantitatively governs how well-trained real-world neural networks separate data throughout all layers. This law has been demonstrated to provide useful guidelines and insights into deep learning practice through network architecture design, training, and interpretation of predictions.

One avenue for future research is to explore the applicability of the law of equi-separation to a wider range of network architectures and applications. For example, it would be interesting to see whether this law holds for neural ordinary differential equations and, if so, how it could be used to improve our understanding of the depth of these continuous-depth models. Additionally, it would be valuable to examine the law in regression settings and multimodal tasks (40). A crucial question that remains to be addressed is understanding why the law is most accurate in feedforward neural networks and becomes somewhat fuzzier in certain convolutional neural networks. Next, given that the law is most evident in feedforward neural networks, an interesting direction is to explore measures in lieu of the separation fuzziness defined in Eq. 1. The hope is that this could potentially make the law clearer in other network architectures. In doing so, it would be beneficial to take into account network structures, such as convolution kernels in convolutional neural networks, during the development of separation measures. Furthermore, while the law of equi-separation has been observed in a variety of vision tasks, it does not appear to hold for language models such as BERT (41) (SI Appendix, Fig. S9). This could be due to the fact that language models learn token-level features rather than sentence-level features at each layer. As such, it would be important to study whether the law of equi-separation could be utilized to derive better sentence-level features from a sequence of token-level features.

Perhaps the most pressing question is to reveal the underlying mechanism of this data-separation law. However, some commonly used assumptions in the literature of deep learning theory may not be applicable. For instance, the nonlinearity of activation functions is crucial to this law because the separation fuzziness is roughly unchanged by linear transformations. Additionally, analyses of neural networks using the neural tangent kernel (42) are not able to capture the effect of depth, which is essential for understanding the data separation process across all layers. This suggests that more innovative approaches may be necessary to fully understand the mechanisms behind this law.

Broadly speaking, a perspective informed by this law is to focus on the nontemporal dynamics of the data separation process throughout all layers when studying deep neural networks. This perspective probes into the interior of deep learning models by perceiving the black-box classifier as a sequence of transformations, each of which corresponds to a layer or a block in the case of residual neural networks. In contrast, existing analyses often consider the “exterior” of these models by focusing on the loss function or test performance, without considering the intermediate layers (42–44). Informed by the law of equi-separation, this perspective may provide formal insights into the inner workings of deep learning.

Supplementary Material

Appendix 01 (PDF)

Acknowledgments

We are grateful to X.Y. Han for very helpful comments and feedback on an early version of the manuscript. We also thank Hangyu Lin (https://github.com/avalonstrel/NeuralCollapse/tree/mlp), Cheng Shi (https://github.com/DaDaCheng/Re-equi-sepa), and Ben Zhou (https://github.com/Slash0BZ/data-separation/tree/main/reproduction) for reproducing our experimental results. This research was supported in part by NSF grants CCF-1934876 and CAREER DMS-1847415 and an Alfred Sloan Research Fellowship.

Author contributions

H.H. and W.J.S. designed research; H.H. and W.J.S. performed research; H.H. and W.J.S. contributed new reagents/analytic tools; H.H. and W.J.S. analyzed data; and H.H. and W.J.S. wrote the paper.

Competing interests

The authors declare no competing interest.

Footnotes

This article is a PNAS Direct Submission.

*For clarification, D0 is calculated from the raw data, and D1 is calculated from the data that have passed through the first layer but not the second layer.

Data, Materials, and Software Availability

Our code is publicly available at GitHub (https://github.com/HornHehhf/Equi-Separation) (45).

Supporting Information

References

  • 1.Krizhevsky A., Sutskever I., Hinton G. E., “ImageNet classification with deep convolutional neural networks” in Advances in Neural Information Processing Systems (Curran Associates, Inc., 2012). vol. 25. [Google Scholar]
  • 2.LeCun Y., Bengio Y., Hinton G., Deep learning. Nature 521, 436–444 (2015). [DOI] [PubMed] [Google Scholar]
  • 3.Silver D., et al. , Mastering the game of go with deep neural networks and tree search. Nature 529, 484–489 (2016). [DOI] [PubMed] [Google Scholar]
  • 4.Fawzi A., et al. , Discovering faster matrix multiplication algorithms with reinforcement learning. Nature 610, 47–53 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Hastie T., Tibshirani R., Friedman J. H., The Elements of Statistical Learning: Data Mining, Inference, and Prediction (Springer, 2009), vol. 2. [Google Scholar]
  • 6.Hutson M., Has artificial intelligence become alchemy? Science 360, 478 (2018). [DOI] [PubMed] [Google Scholar]
  • 7.Bartlett P. L., Long P. M., Lugosi G., Tsigler A., Benign overfitting in linear regression. Proc. Natl. Acad. Sci. U.S.A. 117, 30063–30070 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.Y. Lu, A. Zhong, Q. Li, B. Dong, “Beyond finite layer neural networks: Bridging deep architectures and numerical differential equations” in International Conference on Machine Learning (2018), pp. 3276–3285.
  • 9.Fang C., He H., Long Q., Su W. J., Exploring deep neural networks via layer-peeled model: Minority collapse in imbalanced training. Proc. Natl. Acad. Sci. U.S.A. 118, e2103091118 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.K. He, X. Zhang, S. Ren, J. Sun, “Deep residual learning for image recognition” in 2016 IEEE Conference on Computer Vision and Pattern Recognition (2016), pp. 770–778.
  • 11.D. P. Kingma, J. Ba, “Adam: A method for stochastic optimization” in International Conference on Learning Representations (2015).
  • 12.Xiao H., Rasul K., Vollgraf R., Fashion-MNIST: A novel image dataset for benchmarking machine learning algorithms. arXiv [Preprint] (2017). http://arxiv.org/abs/1708.07747 (Accessed 1 June 2023).
  • 13.Stevens J. P., Applied Multivariate Statistics for the Social Sciences (Routledge, 2012). [Google Scholar]
  • 14.Papyan V., Han X., Donoho D. L., Prevalence of neural collapse during the terminal phase of deep learning training. Proc. Natl. Acad. Sci. U.S.A. 117, 24652–24663 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.X. Han, V. Papyan, D. L. Donoho, “Neural collapse under MSE loss: Proximity to and dynamics on the central path” in International Conference on Learning Representations (2022).
  • 16.K. Simonyan, A. Zisserman, “Very deep convolutional networks for large-scale image recognition” in International Conference on Learning Representations (2015).
  • 17.G. Huang, Z. Liu, L. Van Der Maaten, K. Q. Weinberger, “Densely connected convolutional networks” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (2017), pp. 4700–4708.
  • 18.G. Alain, Y. Bengio, “Understanding intermediate layers using linear classifier probes” in International Conference on Learning Representations (2017).
  • 19.Papyan V., Traces of class/cross-class structure pervade deep learning spectra. J. Mach. Learn. Res. 21, 10197–10260 (2020). [Google Scholar]
  • 20.T. Galanti, A. György, M. Hutter, “On the role of neural collapse in transfer learning” in International Conference on Learning Representations (2022).
  • 21.Galanti T., On the implicit bias towards minimal depth of deep neural networks. arXiv [Preprint] (2022). http://arxiv.org/abs/2202.09028 (Accessed 1 June 2023).
  • 22.Ben-Shaul I., Dekel S., “Nearest class-center simplification through intermediate layers” in Topological, Algebraic and Geometric Learning Workshops 2022 (PMLR, 2022), pp. 37–47. [Google Scholar]
  • 23.Bengio Y., Learning deep architectures for AI. Found. Trends Mach. Learn. 2, 1–127 (2009). [Google Scholar]
  • 24.Schmidhuber J., Deep learning in neural networks: An overview. Neural Networks 61, 85–117 (2015). [DOI] [PubMed] [Google Scholar]
  • 25.Hihi S., Bengio Y., “Hierarchical recurrent neural networks for long-term dependencies” in Advances in Neural Information Processing Systems, Touretzky D., Mozer M., Hasselmo M., Eds. (MIT Press, 1995), vol. 8. [Google Scholar]
  • 26.Glorot X., Bengio Y., Understanding the difficulty of training deep feedforward neural networks. Statistics 9, 249–256 (2010). [Google Scholar]
  • 27.Yarotsky D., Error bounds for approximations with deep ReLU networks. Neural Networks 94, 103–114 (2017). [DOI] [PubMed] [Google Scholar]
  • 28.LeCun Y., The MNIST database of handwritten digits (1998) http://yann.lecun.com/exdb/mnist/. Accessed 1 June 2023.
  • 29.A. Krizhevsky, “Learning Multiple Layers of Features from Tiny Images,” Master’s thesis, University of Toronto (2009).
  • 30.K. He, J. Sun, “Convolutional neural networks at constrained time cost” in 2015 IEEE Conference on Computer Vision and Pattern Recognition (2015), pp. 5353–5360.
  • 31.Srivastava R. K., Greff K., Schmidhuber J., “Training very deep networks” in Advances in Neural Information Processing Systems (Curran Associates, Inc., 2015). vol. 28. [Google Scholar]
  • 32.Howard A. G., et al. , Mobilenets: Efficient convolutional neural networks for mobile vision applications. arXiv [Preprint] (2017). http://arxiv.org/abs/1704.04861 (Accessed 1 June 2023).
  • 33.Tan M., Le Q., EfficientNet: Rethinking model scaling for convolutional neural networks. Learning 97, 6105–6114 (2019). [Google Scholar]
  • 34.Bubeck S., Sellke M., “A universal law of robustness via isoperimetry” in Advances in Neural Information Processing Systems (2021), vol. 34, pp. 28811–28822. [Google Scholar]
  • 35.Ioffe S., Szegedy C., Batch normalization: Accelerating deep network training by reducing internal covariate shift. Learning 37, 448–456 (2015). [Google Scholar]
  • 36.Zhang C., Bengio S., Singer Y., Are all layers created equal? J. Mach. Learn. Res. 23, 1–28 (2022). [Google Scholar]
  • 37.Su W. J., Neurashed: A phenomenological model for imitating deep learning training. arXiv [Preprint] (2021). http://arxiv.org/abs/2112.09741 (Accessed 1 June 2023).
  • 38.M. D. Zeiler, R. Fergus, “Visualizing and understanding convolutional networks” in European Conference on Computer Vision (2014), pp. 818–833.
  • 39.I. Tenney, D. Das, E. Pavlick, “BERT rediscovers the classical NLP pipeline” in Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics (2019), pp. 4593–4601.
  • 40.Liang V. W., Zhang Y., Kwon Y., Yeung S., Zou J. Y., “Mind the gap: Understanding the modality gap in multi-modal contrastive representation learning” in Advances in Neural Information Processing Systems (2022), vol. 35, pp. 17612–17625. [Google Scholar]
  • 41.Devlin J., Chang M. W., Lee K., Toutanova K., BERT: Pre-training of deep bidirectional transformers for language understanding. Long Short Pap. 1, 4171–4186 (2019). [Google Scholar]
  • 42.A. Jacot, F. Gabriel, C. Hongler, “Neural tangent kernel: Convergence and generalization in neural networks” in Advances in Neural Information Processing Systems (2018), vol. 31.
  • 43.H. He, W. J. Su, “The local elasticity of neural networks” in International Conference on Learning Representations (2020).
  • 44.Belkin M., Hsu D., Ma S., Mandal S., Reconciling modern machine-learning practice and the classical bias-variance trade-off. Proc. Natl. Acad. Sci. U.S.A. 116, 15849–15854 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 45.He H., Equi-Separation. GitHub. https://github.com/HornHehhf/Equi-Separation. Deposited 1 May 2023.

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Appendix 01 (PDF)

Data Availability Statement

Our code is publicly available at GitHub (https://github.com/HornHehhf/Equi-Separation) (45).


Articles from Proceedings of the National Academy of Sciences of the United States of America are provided here courtesy of National Academy of Sciences

RESOURCES