Abstract
The architecture of early olfactory processing is a striking example of convergent evolution. Typically, a panel of broadly tuned receptors is selectively expressed in sensory neurons (each neuron expressing only one receptor), and each glomerulus receives projections from just one neuron type. Taken together, these three motifs—broad receptors, selective expression, and glomerular convergence—constitute “canonical olfaction,” since a number of model organisms including mice and flies exhibit these features. The emergence of this distinctive architecture across evolutionary lineages suggests that it may be optimized for information processing, an idea known as efficient coding. In this work, we show that by maximizing mutual information one layer at a time, efficient coding recovers several features of canonical olfactory processing under realistic biophysical assumptions. We also explore the settings in which noncanonical olfaction may be advantageous. Along the way, we make several predictions relating olfactory circuits to features of receptor families and the olfactory environment.
I. INTRODUCTION
Chemosensation is our oldest sensory modality, and for most animals it is the primary means of sensing the environment. The chemoreceptors underlying this process have a rich evolutionary history which mirrors the specialization for diverse habitats across species [1–5]. Chemoreceptors evolve with dizzying speed, evincing their position at the interface between the organism and an ever-changing chemical landscape [6, 7]. Surprisingly, however, the olfactory circuit in which these receptors are embedded has deep similarities across vertebrates and invertebrates [8, 9]. In both lineages, organisms have evolved a large repertoire of receptors, many of them broadly tuned (though some receptors are specialists for a given ligand of high importance) [10, 11]. Each primary sensory neuron typically expresses just one receptor, and only neurons expressing the same receptor converge onto a given olfactory glomerulus.
The transcriptional and wiring mechanisms for achieving this circuit organization vary widely across animals. This diversity suggests that these three motifs (broadly-tuned receptors, one neuron-one receptor, and glomerular convergence) are the result of strong selective pressure, rather than evolutionary chance or molecular constraints. Why might this architecture be optimal?
To answer this question, we must specify an objective. Organisms rely on olfaction for many of their basic needs, including mating, feeding, and avoiding predation. From a computational perspective, this means that animals need to solve a wide array of tasks, including discriminating odor identity [11–13], segmenting odor landscapes [14–18], and matching single odorants to a known mixture (“pattern completion”) [19]. Complicating the picture further are the rich temporal dynamics of olfactory stimuli, which induce correspondingly rich dynamics even in the earliest levels of sensory processing [14, 20, 21].
In light of this complexity, we consider a task-agnostic objective: maximizing mutual information. The idea that early sensory processing is organized to maximize mutual information between stimulus and neural representation is known as efficient coding [22, 23]. Of course it is true that in practice, animals benefit from discarding irrelevant information. But critically, the relevance of information is often context- and task-dependent. Since higher areas of the animal brain can only access the olfactory stimulus through the glomeruli, it seems likely that up to the glomerular layer, information maximization is a plausible approximate objective. Many previous works have applied efficient coding principles to olfaction under this reasoning [24–29].
In this paper, we model olfactory stimuli and receptor activity with widely-used existing models that capture the relevant biophysics while remaining computationally tractable [24, 26]. To allow tractable analysis, we ignore temporal response dynamics in this work (but see the Discussion). Within this framework, we leverage recent technical advances from the domain of unsupervised representation learning to numerically maximize a proxy for the mutual information between the stimulus and the corresponding neural response. We do this with respect to three matrices: the receptor by odorant sensing matrix W, the neuron by receptor expression matrix E, and the glomerulus by neuron connectivity matrix G. In order to better respect the different timescales at play in the evolution and development of the circuit, we optimize one layer at a time. We find that efficient coding recovers the motifs of olfactory architecture that have now been described in several species separated by hundreds of millions of years of evolutionary time [9, 30–33]. Thus, this striking example of convergent evolution can be understood through the lens of optimality for information processing.
II. SETUP
A. Stimulus and encoding model
We model the olfactory stimulus as a sparse vector in the space of concentrations of N monomolecular odorants (Figure 1a). We first generate a binary vector using a Gaussian copula. This allows us to tune the mean and covariance of . Means are drawn from a Gamma distribution, so that some odorants are frequent, but most are rare (Figure 1b). We set the covariance of to be a block matrix, where each of the k blocks represents a mixture emitting correlated odorants (Figure 1c). A more realistic covariance matrix would have a nested, hierarchical structure rather than merely a single level of blocks [34], but it is less obvious how to parametrize such matrices—see Section IIIB and the Discussion.
FIG. 1.
Overview of the model. (a) The structure of the stimulus and the network. (b) The frequency of each odorant across samples. (c) An example covariance matrix of the binarized odorants. The number of blocks (each representing a source) is a parameter we vary—see Section III B.
We then assign an independent and identically distributed log-normal concentration with variance to each odorant in the sample to generate . A version of this sparse, log-normal model for was used by Qin et al. [24], and we extend their model by incorporating varying frequencies, structured covariance, and a degree of sparsity that changes across samples. We include full details of the statistical model in the Methods.
Note that while we preserve the presence/absence statistics across samples (so that odorant α can occur more frequently than odorant β), we do not preserve odorant concentration ratios across samples (odorant α cannot occur at a typically higher concentration than odorant β). Such differences in concentration can be a critical aspect of identity coding [35]. However, this structure is liable to disruption by the turbulent transport of volatile molecules to the sensory epithelium, as well as differences in molecular diffusivity [35, 36]. Neglecting this structure greatly reduces the number of tunable parameters in our stimulus model.
Given the stimulus, the encoding model is
| (1) |
where is the firing rate of the neurons, is the receptor by odorant affinity matrix, is the neuron by receptor expression matrix, is a Hill function with coefficient n, applied element-wise, and . This form of nonlinearity is consistent with measurements in the fly larva by Si et al. [37], in which the authors found that a Hill coefficient of best fit their data. The coefficient is typically higher in vertebrates (on average ), where olfactory receptors are G protein-coupled receptors rather than ligand-gated ion channels [38, 39] (but see Grosmaitre et al. [40] for in vitro). We confirmed that the qualitative features of the optimal W and the optimal E did not depend on this coefficient (SI Appendix: Fig. S1). We discuss limitations of this model, as well as possible improvements, in the Discussion section. While both W [24, 26] and E [25] have been studied using efficient coding, most works have not considered the interplay between the two (with the exception of Lienkaemper et al. [41]).
Finally, the neural activity is integrated in regions of neuropil known as glomeruli , with . For simplicity, we match the number of glomeruli to the number of receptors, although in reality there could be more glomeruli (as in mice [42]) or fewer (as in mosquitoes [43]). Here the vector g consists of one representative dendrite per glomerulus. In paired recordings from pre- and post-synaptic neurons at the glomerulus in Drosophila melanogaster, Bhandawat et al. [44] found that the transfer function was sigmoid in shape, so we parametrize it using the hyperbolic tangent with varying gain α:
| (2) |
B. Layer-wise efficient coding
Given this model, we perform three separate optimizations. The first is over receptor sensitivities:
| (3) |
By , we indicate that we are plugging in a canonical (one neuron-one receptor) E matrix into (1) to optimize over W. Thus we are modeling the ongoing evolution of receptors in the context of the widespread one neuron-one receptor motif, rather than the joint emergence of the two phenomena. When optimizing W, we constrain its elements to be positive (see Methods for details of the optimization procedures for all three layers).
The next optimization we perform is
| (4) |
in which we first optimize W as in (3), then optimize over E. We also shuffle W to ensure that our conclusions are robust to the details of the receptor matrix. Varying the details of W and the parameters of the environment allows us to probe why canonical expression is so common, as well as exploring the phenomenon of noncanonical expression, as recently characterized in the Aedes aegypti mosquito [43, 45]. We constrain E to be positive and to sum row-wise to unity, since neurons must allocate a finite budget of receptor expression [43, 45–48].
Finally, we optimize the glomerular layer:
| (5) |
where is a shuffled version of . This is likely to be a more realistic model of true receptors than due to the “sloppy” tuning of real receptors (something we discuss at length in Section IIIB). Again we constrain G to be positive and to sum row-wise to unity, since glomeruli receive excitatory inputs from finitely many olfactory receptor neurons [8].
We emphasize that we do not jointly optimize over all three matrices, but rather treat one layer at a time, starting with W. We consider the trade-offs of this layer-wise approach, its alternatives, and its biological interpretation in the Discussion.
C. Maximizing mutual information by proxy
The primary challenge in any efficient coding analysis is the estimation of mutual information. To enable analytical progress, strong assumptions on the stimulus and encoding are required [24, 25, 41, 49–51]. Most simply, if one assumes linear processing of a Gaussian stimulus, the mutual information can be computed in terms of the log determinants of the resulting covariance matrices. However, it is challenging to analytically compute mutual information in more realistic non-Gaussian models like the one we adopt here.
Numerical estimation of mutual information in high dimensions is also fraught with difficulties [50, 52]. Here, inspired by work on deep representation learning, we adopt a conservative approach based on the principle that finding an optimal encoding model does not require an estimate of the precise value of the mutual information. Instead, it is sufficient to maximize a reliably-estimable proxy objective function [53, 54]. Committing to this approach limits our ability to compare different optimized models, but it allows for a much more robust and efficient optimization procedure.
As detailed in the Methods, we leverage a variational formulation for mutual information maximization first proposed by Nowozin et al. [55]. Specifically, we maximize a bound on the Jensen-Shannon divergence (JSD) between the joint and marginal distributions of the stimulus and the encoding. This was inspired by previous works showing that the JSD can be estimated more reliably than the Kullback–Leibler divergence (KLD) that defines the mutual information [53, 54]. We describe the trade-offs of this approach in the Methods. We handle constraints on W, E, and G using the framework of mirror descent; see Methods for details.
III. RESULTS
A. Receptors
We first optimized the mutual information over the sensing matrix W. Each row of this matrix represents a receptor, and each column an odorant. Thus is the sensitivity of the i-th receptor to the j-th odorant. It is helpful to consider that chemoreceptors are both ancient (predating neurons, for example), and fast to evolve—indeed, they are some of the fastest evolving proteins in many organisms [2, 6, 7, 56]. This may be because their conserved function is less critical for survival than that of metabolic or structural proteins. Furthermore, there is very direct evolutionary pressure to modify chemoreceptors in response to changes in the organism’s chemical environment.
Thus it seems that the matrix W is relatively tunable (notwithstanding biophysical constraints—see Discussion), and so we asked: to what extent can efficient coding recover the qualitative features of receptor affinities measured in biological circuits? To guide our analysis, we turned to published data on the well-characterized olfactory system of the Drosophila melanogaster larva, which has just 21 olfactory receptor neurons, each identified by the expression of a unique receptor. The tractable size of this system permits comprehensive interrogation of the complete receptor repertoire, which was conducted by Si et al. [37]. Plugging in a canonical expression pattern E as in equation (3), we optimized mutual information between neural activity r and the stimulus c, using parameters that matched the fly larva circuit.
The resulting had a non-negligible degree of sparsity (≈ 40% entries = 0; see discussion in SI Appendix A), and its non-zero elements were well-fit by a log-normal distribution—see SI Appendix: Figs. S2 and S3. Both of these features have been derived analytically in previous theoretical work [24, 26] in the limit of vanishing neural noise. Here we confirmed that these results still approximately hold when neural noise is non-negligible. We next sought to compare our optimized W to experimental results in a more detailed and biology-focused analysis, as shown in Figure 2.
FIG. 2.
Optimizing over W for parameters typical of the fly larva, given canonical expression. (a) The trajectory of mutual information over the course of the optimization. (b) Sensitivity per odorant in the optimization (blue) and fly larva (red). Each tick denotes the sensitivity of one receptor for that odorant. In both cases, there are typically many receptors per odorant, and their sensitivities span several orders of magnitude. Odorants are sorted by their maximum sensitivity within each panel. (c) Distribution of . For the optimizations, 20 runs with different random seeds are shown. (d) More frequent odorants are afforded greater dynamic range. Error bars denote standard deviation over 20 runs (only the upper error bars are shown). (e) For fixed , the number of “specialist receptors” grows with the number of receptors. “Fit” is an analytic control obtained by fitting a log normal distribution to the optimized W matrix. Experimental data in panels (b) and (c) is from Si et al. [37].
We found three qualitative matches between our optimized W and the fly larva W. First, the distribution of sensitivities across all receptors and odorants is broad, spanning approximately six orders of magnitude (see Figure 2c). This heavy-tailed distribution is the object of much analysis by Si et al. [37], who argue for the computational advantages of such a code. In our results, it is difficult to confirm a precise power law to the exclusion of other heavy-tailed distributions. But the qualitative agreement shown in Figure 2c may further support the optimality of this distribution for the entries of W.
Second, each odorant typically has at least several receptors tuned to it, as shown in Figure 2b. This is necessary to transmit faithful information about the stimulus, since the two orders of magnitude spanned by each receptor’s Hill function activity do not cover the full range of odorant concentrations. (For Figure 2, odorant concentrations were drawn from a log-normal distribution with σc = 3.)
Third, some odorants are sensed by a greater number and range of receptors than others (compare the top and bottom lines of Figure 2b). Within the context of our optimization, there is an intuitive reason for this: it is the more frequent odorants that are afforded greater dynamic range, as shown in Figure 2d. While this result may hold in experimentally measured receptor sensitivities, testing it directly would require quantitative knowledge about natural olfactory landscapes, which are notoriously difficult to characterize.
More fundamentally, however, this last finding underscores one limitation of the efficient coding framework. Mutual information is a statistical measure of the dependence between two distributions and does not privilege any particular component of the stimulus for reasons of ecology or behavior. Thus, when optimized, our model circuit simply devotes greater resources to more common odorants. In natural settings, by contrast, animals may have evolved receptors to sense an infrequent but critically important odorant (such as a toxin or pheromone) with high sensitivity across a wide range of concentrations, as in Sakurai et al. [57]. Thus, although the trend in Figure 2d may hold on average, it is likely to admit important exceptions in biological circuits.
The latter observations consider the circuit’s ability to report information about a given odorant (studying columns of W). The complementary viewpoint is to begin with a receptor and study its tuning across odorants (rows of W). For example, much of the recent experimental work on olfactory receptors has emphasized the promiscuity of their binding. Rather than a “lock and key” pairing, ligands were found to fit loosely in the binding pocket of the insect receptor MhOR5 [58]. This strategy is the natural result of the odorant to receptor bottleneck (), and under our model, most receptors indeed exhibit broad tuning (SI Appendix: Fig. S2). However, this bottleneck argument does not explain why some receptors seem to have narrow tuning for one or a handful of odorants [10], especially in organisms with greater numbers of receptors (such as in mice, where M ≈ 1100).
To explore this question, we varied the degree of the bottleneck (keeping N fixed and increasing M) and inspected the resulting optimized receptor profiles. As we increased the number of receptors, the circuit acquired an increasing number of “specialist receptors” (see Figure 2e). We formalized this using the following criterion: if the maximum sensitivity of a receptor for an odorant was at least two orders of magnitude greater than the 99th percentile sensitivity, then we counted it as a specialist for that odorant. This corresponds to testing a panel of 100 odorants, and finding that a receptor has 100x greater sensitivity for its specialized odorant than for all other odorants in the panel.
In Figure 2e, we plot two important controls. One is a shuffled W. For this control, the trend still holds, which indicates that specialist receptors are developed by pushing up the sensitivity of the specialist receptorligand pair, rather than pushing down the sensitivities of the specialist receptor for other ligands. Therefore, a specialist in our model is a receptor with mostly typical sensitivities, and one very high outlier.
Our other control is an analytic fit. Here we fit a log-normal distribution to the values for each value of M, then sample from that distribution to generate a W, and count specialists in that W. While the log-normal is heavy-tailed and fits well for the bulk of the distribution (SI Appendix: Fig. S3), it does not generate the extreme outliers needed to produce any specialist receptors under our very stringent criterion. This demonstrates that the effect is not driven simply by drawing more samples from the same distribution.
These findings suggest that specialist receptors may “look normal” until their target ligand is found. Conversely, given knowledge about a specialist receptor, it may be worth testing other odorants against that receptor to see if they are detected at reasonable concentrations (such as in Meyerhof et al. [59]).
Interestingly, despite the allocation of greater dynamic range to higher frequency odorants (Figure 2d), we did not find that they were more likely to be targeted by specialized receptors (SI Appendix: Fig. S4). This may reflect the fact that, when specialist receptors are saturated in the presence of their target odorant, they effectively exacerbate the bottleneck faced by the rest of the circuit . This suggests that there may be slight pressure against developing specialist receptors for common odorants.
In this way, optimal sensing matrices W may leverage structure in the stimulus, as well as promiscuous binding, to get useful information from a specialist receptor when the target odorant is not present. Some support for this idea can be found in studies [60, 61] which showed that background odors can compromise the detection of a ligand by its cognate specialist receptor. Conversely, host plant volatiles were shown to synergistically improve pheromone detection in the silkmoth Bombyx mori at the receptor neuron level [62].
Finally, one important check on our numerical procedure was to confirm that the distribution of the optimal W does not depend on initialization. In Figure 2, the initial W is a scaled log-normal: , where is the mean magnitude of the stimulus vector (see Methods). We confirmed that scaling the log-normal to the minimum sensitivity instead, below which is effectively 0, did not change the results in Figure 2. (We derive this value and discuss its implications in SI Appendix A.)
Initialization had virtually no effect on the distribution of (as in panels b,c) or the relationship between frequency and dynamic range (panel d). The one instance of dependence on initialization was seen in the high M regime () in panel e. With so many receptors, the bulk of could remain close to initialization without degrading performance. This increased the number of specialists that emerged by roughly a factor of 2 (maximum 57 rather than 27 as in Figure 2), but did not change the trend. We plot Figure 2 for this alternative initialization in the SI Appendix: Fig. S5.
To summarize, the which results from our optimization procedure shares a number of properties with the experimental W measured in the fly larva. The distribution of values spans six orders of magnitude and is well fit by a log-normal density. Receptor tuning is surprisingly broad, given that we do not explicitly model the imperfect binding of real ligand-receptor pairs. Lastly, increasing M permits the emergence of specialists in circuits with larger receptor families.
B. Expression
Perhaps the most striking feature of canonical olfaction is the “one neuron-one receptor” rule. From a computational perspective, it is not immediately clear why this should be optimal, and yet it has emerged in various organisms including flies [31], mice [32], and ants [33]. Notably, each of these organisms has developed an entirely different mechanism to generate this pattern of expression.
In flies, a core set of 15-20 transcription factors and cis-regulatory elements acts combinatorially to establish the expression of a single odorant receptor (OR) in a given cell type [63]. In mice, on the other hand, the OR loci from 18 chromosomes come together in a hub that chooses a single receptor allele for expression and silences all the rest [32]. Very recent work has demonstrated how a complex gene expression program whose usage varies tightly with position in the olfactory epithelium orchestrates this OR choice [47, 48]. In ants, many odorant receptors are transcribed, but only the most upstream gene is exported out of the nucleus [33, 64].
These and other disparate mechanisms [63] suggest that the one neuron-one receptor rule may confer a strong fitness advantage. Accordingly, when we numerically maximized our proxy mutual information with respect to E in (1) for numbers typical of the fly larva, we found that a clear pattern of canonical expression emerged, as shown in Figure 3. This was despite random initialization (Figure 3b).
FIG. 3.
Optimization over E for numbers typical of the fly larva. (a) The trajectory of mutual information over the course of the optimization. (b) The initial expression matrix. (c) The optimized expression matrix. The optimization is over E in (1), plugging in the optimal W from Figure 2. (d) The fraction of variance in neural activity r explained by each of the top principal components, given the initial expression from panel (b) and the optimized expression from panel (c). Statistics are computed over 1000 samples.
Immediately, however, we are confronted with a “chicken-or-egg” problem. In (3), we optimize W given canonical E, then in (4) we optimize E with the resulting . This could in principle bias the optimization over E to converge on canonical expression. To account for this possibility, in our subsequent analysis we optimize E given five plausible alternative models for W. We find that each of these models favors canonical expression (see analysis below, and SI Appendix: Fig. S11), which indicates that the result is not merely an artifact of our layer-wise approach.
Given that the solution is robust, then why is canonical expression optimal? In the limit of low neural noise, one simple way to increase mutual information is simply to decorrelate activity, when a distribution of activity is computed over the stimulus distribution p(c). This is a standard prediction of many efficient coding analyses [23, 65]. Accordingly, we find that the distribution of variances explained by the principal components of activity is flatter under canonical expression than under random expression (Figure 3d). The best way to decorrelate activity in this way is to place neurons at the corners of the simplex in gene expression space, since total receptor expression per neuron must sum to unity. This corresponds precisely to the one neuron-one receptor rule. (See Lienkaemper et al. [41] for a related analysis in the case of one neuron.)
We next sought to understand if the same result would hold in a larger system. We chose a scale comparable to the adult fly, with M = 60 receptors and L = 1260 olfactory neurons. This is a natural focus, since olfactory processing in the fly is well characterized (see for example Hallem and Carlson [11]), and these numbers permit reasonable numerics. We first optimized W, and found a qualitative match to the measurements in Hallem and Carlson [11], with most receptors exhibiting broad tuning and a handful of receptors exhibiting narrow tuning (SI Appendix: Fig. S6).
Surprisingly, when we optimized over E using this W matrix, a sloppier expression matrix emerged, in which a handful of receptors were typically coexpressed per neuron, as shown in Figure 4a. Interestingly, recent work in the Aedes aegypti mosquito has discovered this qualitative pattern of expression [43, 45]. Reanalysis of existing data by Adavi et al. [45] further characterized occasional exceptions to the one neuron-one receptor rule in Drosophila. The extent of this coexpression is controversial [63, 66] but some unambiguous cases have been known for many years [67, 68].
FIG. 4.
Optimization over for numbers typical of adult Drosophila. (a) The optimized expression obtained from plugging in the optimized sensing matrix . (b) The optimized obtained from plugging in a shuffled version of . (c) Covariance between receptor pairs which are coexpressed in at least three neurons (“robust coexpression”, purple), receptor pairs which are singly expressed (blue), and receptor pairs (orange). Covariances are computed on log sensitivities. (d) Probability of robust coexpression given binned covariances. Data in (c) and (d) are across 21 runs with varying environmental parameters and , given by the bottom three rows of the phase diagram in Figure 5. The p-value for the difference in distribution between robustly coexpressed pairs (purple, ) and singly expressed pairs (blue, ) is obtained using the Mann-Whitney U test.
As we interpret this result, it is is important to note that, in both Aedes aegypti and Drosophila, the majority of coexpressed chemoreceptors are disproportionately close to each other in genomic and phylogenetic space [45]. Adavi et al. [45] term this “coexpression by descent.” In only a minority of cases did the authors find that coexpressed receptors were far from each other (“coexpression by co-option.”) This suggests that much of the phenomenon may be a by-product of shared gene regulation. Since families of chemoreceptors frequently expand via tandem duplications (the “birth and death model” of gene evolution) [7, 69], there are many examples of neighboring ORs with similar sequences whose coexpression might be driven by cis-regulatory elements [63].
These molecular mechanisms likely have little to do with optimal information processing. However, the fact remains that different insects (like Drosophila and Aedes aegypti) exhibit strikingly different levels of coexpression. Our results suggest that similarity between receptor affinity profiles is one factor in understanding the functional significance of these differences, as shown in Figure 4c and d. Since phylogenetically related receptors typically have similar affinity profiles [70, 71], this is consistent with the analysis of Adavi et al. [45]. Under this reasoning, coexpression is primarily driven by shared regulatory machinery, but once it emerges, its fitness effect depends on the similarity of the coexpressed receptors, with coexpression of more similar receptors being less detrimental.
Given these subtleties, why did we obtain such a clean canonical expression profile in our analysis of the larval fly (Figure 3)? The answer may be in the number of receptors M. Sweeping over this parameter, we found that receptor matrices with small M were much closer to full rank (SI Appendix: Fig. S7), indicating that their receptors were more differentiated. It may be possible that organisms with only a handful of evolutionarily mature receptors cannot afford similar affinity profiles. Conversely, circuits with many receptors are “over-parameterized” and can find many solutions to the problem, coexpressing without penalty.
To ensure that these results did not depend on specific parameter values, we next swept over the two important parameters of our odorant model: the structure of the covariance matrix and the log-normal noise parameter . We parameterized by giving it a block structure and varying the number of blocks (see Figure 1 and the Methods). Ecologically, each block of correlated odorants corresponds to a source in the organism’s environment. Precise quantification of the fly’s natural olfactory environment is elusive (see Yang et al. [72] and Zhou et al. [34] for recent efforts in this direction), but it is plausibly contained in the large space we explore.
One trend in the resulting phase diagram (Figure 5a) is that limited noncanonical olfaction is favored for increasing environmental noise . This agrees with very recent work by Lienkaemper et al. [41], who approach the problem using a linear-Gaussian theory. In further accord with their analysis, we find that having fewer sources, each emitting more correlated odorants, supports more noncanonical olfaction, at least for low levels of environmental noise (see bottom two rows of Figure 5a). The magnitudes of these effects, however, are small (Figure 5c, d). Note that our model is not directly comparable with theirs, due to differences in the formulation of stimulus versus noise (see Discussion).
FIG. 5.
The degree of canonical expression in optimized matrices while varying environmental parameters and . (a) The phase diagram using optimized . (b) The phase diagram using shuffled . (c) through (f): example expression matrices taken from the above phase diagrams. The score is , where measures the entropy of the -th row of , when the expression is viewed as a probability distribution over receptors, and the expectation is taken over the rows of .
Since our primary focus is on understanding why canonical olfaction is so common, in Figure 5 we have initialized expression matrices to be noncanonical: each neuron expresses 3-7 receptors at roughly equal levels, for a “canonical score” of 0.74 (see Figure 5 caption for definition). This was done in order to determine the degree to which canonical expression is strongly favored compared to receptor coexpression. It is clear that most of the optimizations in Figure 5a barely move, indicating that this noncanonical pattern at initialization is already locally optimal. However, when we initialize expression to be canonical, we also see virtually no change (SI Appendix: Fig. S8). Therefore, neither canonical nor noncanonical expression is strongly favored, regardless of environmental structure. Mathematically, both expression patterns represent solutions to the optimization.
To better understand this puzzling result, we inspected the optimized matrices . We found a distinctive low rank structure in across environmental parameters that was especially pronounced for higher values of (SI Appendix: Fig. S9). This is driven by our frequency model since, as indicated in Figure 1b, only a small subset of odorants occur with high frequencies, while the bulk occur at low frequencies. Flattening the frequency distribution abolished the low rank structure (SI Appendix: Fig. S10). This degeneracy in permits either canonical or noncanonical olfaction, since the precise details of expression do not matter when W is low rank. Importantly, this finding is likely to depend on the details of our biophysical model (1), especially the placement of the nonlinearity (see Discussion).
Our modeling is therefore consistent with the hypothesis that if organisms need to sense only relatively few odorants in their environment, then their receptors may share much of their tuning for those odorants. This in turn permits both canonical and noncanonical expression without strong pressure in either direction.
While potentially informative, this result also reflects a limitation of our mutual information objective. In our results, receptor tuning is almost uniformly high for the highest frequency odorants (SI Appendix: Fig. S6), whereas in reality, organisms may need to detect common and rare odorants with equal sensitivity. Additionally, biophysical constraints limit the number of ligands a receptor can bind. Indeed, Si et al. [37] found that covariance in olfactory sensory neuron activity is primarily driven by the geometric structure of odorants. This geometry constitutes a significant constraint on receptor affinities that is not accounted for in our model, and thus real receptors are very likely to be sloppier and less finely tuned than our matrices.
In light of these limitations, we next studied expression using shuffled versions of the optimized W matrices, as in Figure 4b. Shuffling allowed us to preserve the scale and distribution of , while breaking the low rank structure. It also accounts, albeit imperfectly, for the suboptimal tuning of real receptors. Using , we found that canonical expression emerged robustly across almost all combinations of environmental parameters (as shown in Figure 5b), despite noncanonical initialization. Again, we found only a weak dependence on environmental parameters, and the trends were reversed compared to Figure 5a. We tested two other unstructured models for W: shuffling within rows (in order to preserve mean receptor tuning), and fitting a log normal distribution to W. Both of these favored canonical expression when E was then optimized (see SI Appendix: Fig. S11).
Real W matrices are most likely best modeled by something between and . They have some correlation structure, due to the development of new odorant receptors from existing ones, but are unlikely to be extremely low rank. We therefore implemented two more models for W: a log normal analytic W with block co-variance structure, to model families of related receptors, and a log normal W with Toeplitz covariance structure, to model a range of similarities across receptors. We initialized expression to simulate coexpression by descent; that is, coexpression of correlated receptors. However, we found that canonical expression was still largely favored regardless of receptor details (see SI Appendix: Fig. S11).
Finally, we sought to understand how the level of neural noise affected these results. We swept across for a representative pair of environmental parameters (blocks = 64, ) (SI Appendix: Fig. S12). We found that at extremely low noise levels (), both solutions were equally favored, but for more realistic levels (), canonical olfaction was favored as shown above. All of the above results were run for , a value we chose because it sets the noise to be on the order of the mean activity (see SI Appendix: Fig. S13 and Glomerular convergence). Increasing beyond this point further favored canonical olfaction. This trend is consistent with the theory of Lienkaemper et al. [41].
Taken together, these results suggest that receptor co-tuning, rather than environmental statistics, is decisive in determining optimal expression. In our model, only extremely low-rank W matrices permit significant levels of noncanonical olfaction. Biologically, such a set of receptors would have highly correlated tuning across many odorants. A number of more plausible models for W support largely canonical olfaction for realistic levels of neural and environmental noise. This may explain why the one neuron-one receptor rule has emerged in such distantly related organisms, despite the fact that these organisms sense different chemical environments with receptor families that are accordingly divergent.
C. Glomerular convergence
Olfactory sensory neurons converge onto regions of neuropil known as glomeruli [73]. In the canonical model, only neurons expressing the same receptor converge onto a given glomerulus [74]. This “glomerular convergence” confers an intuitive advantage: given canonical expression, the signals from each receptor can be averaged across neurons without mixing across receptors. We plugged in a shuffled version of , the corresponding , and optimized our proxy for mutual information over G and as in (6). We confirmed that glomerular convergence is recapitulated in our model across a wide range of environmental parameters (see Figure 6a–c, and SI Appendix: Fig. S14 for sweep).
FIG. 6.
An example optimization over the glomerular layer for 64 sources and . (a) The course of the mutual information estimate over the optimization. (b) The initial random connectivity. (c) The optimized connectivity. Neurons are sorted by their highest expressed receptor. (d) The transformation at the glomerular layer. A higher serves to flatten the distribution of activity at the glomerulus.
But glomeruli do not just serve to denoise olfactory neuron responses. Careful recording of both the pre-synaptic olfactory receptor neuron (ORN) and the post-synaptic projection neuron (PN) in Drosophila has characterized other aspects of the transformation [44]. Chief among these is “histogram equalization,” in which low ORN activity is amplified but higher activity is not. Since ORN activity is clustered near 0, this induces a more balanced distribution of activity in the post-synaptic neuron, which in turn enhances information transmission. This is a classical prediction from the early days of efficient coding [75].
We first confirmed that our activity clustered near 0 as in Bhandawat et al. [44] (see Figure 6d, bottom axis, and SI Appendix: Fig. S13 for sweep). This is likely necessary due to the log-normally distributed stimulus; if the optimal W is to account for the rare, highest concentration odorants, then the bulk of stimuli will generate responses closer to baseline. Accordingly, we found that as we increased , the distribution of activity was more sharply peaked around 0 (SI Appendix: Fig. S13).
Optimizing over the gain parameter in (5) resulted in a qualitative match to data from adult Drosophila, in which the post-synaptic activity was more evenly distributed across the dynamic range of the neuron (Figure 6d). The resulting histograms were not perfectly equalized, but this may reflect the fact that increasing the entropy of the representation addresses just one term out of two in the expression for mutual information:
| (6) |
In sum, our results are consistent with the idea that glomerular convergence allows for denoising and histogram equalization of olfactory receptor neuron responses when expression is canonical. As we previously observed that canonical expression is largely favored across a range of environmental parameters, we did not investigate what convergence patterns emerge when multiple receptors are expressed in each sensory neuron. However, as long as neural noise is a limiting factor, convergence is the likely solution for denoising.
IV. DISCUSSION
In this work, we set out to understand why odor coding exhibits deep similarities across vertebrates and invertebrates. Using efficient coding, we recovered three motifs of a widely shared olfactory logic: broad receptors, single receptor expression, and glomerular convergence.
A. Receptors
The non-zero sparsity of the optimal W and the broad distribution of its values have been derived in previous theoretical works that approximate the mutual information as the entropy , an approximation which becomes exact in the limit of no neural noise [24, 26]. Our analysis first recovers these results in the regime where neural noise is non-negligible, then builds on them by studying more granular features of such as receptor tuning, odorant sensitivity across receptors, and specialists. These details allow us to make new predictions for experiment, particularly regarding specialist receptors.
In our model, receptor tuning is much more promiscuous (most receptors respond to most odorants) than strictly necessary to map all N odorants in an approximate labeled-line scheme. This correspondence was surprising, since broadly tuned receptors in biology might simply be a byproduct of ligand-receptor biophysics. If this were the case, then such receptors would not emerge in our unconstrained model when we initialize to zero.
Our results suggest instead that broad tuning confers an information processing advantage. The rich theory of compressed sensing may shed light on this advantage, but it is difficult to apply classical results from this field directly to the problem at hand, since a linear measurement model is usually assumed [76]. Recent progress in nonlinear compressed sensing may enable theoretical understanding of the bottleneck problem in olfaction, but much remains unknown [24, 27, 28, 77, 78].
These broadly tuned receptors enable a combinatorial coding scheme in which many more odorants can be represented than there are receptors [79]. To the extent that specialists do emerge in our optimizations, they have increased their sensitivity for the target ligand rather than decreasing their sensitivity for off-target ligands, and are thus still available for sensing other odorants.
One prediction of our analysis is therefore that, in the absence of the target ligand, specialist receptors may be co-opted for the sensing of other relevant molecules at realistic concentrations. This idea is indirectly supported by their diminished performance in the presence of background odorants [60, 61], which indicates that specialist receptors preserve some tuning for off-target ligands. Our results suggest that this phenomenon may be a feature, not a bug, when the specialized ligand is not present. This could be further explored simply by presenting specialist receptors with a typical panel of common odorants, ideally ones not naturally co-occuring with the target ligand.
The question of the relevant range of concentrations for odorant stimuli is critical. Wachowiak et al. [80] have recently argued that the concentrations used in typical experimental studies may exceed those encountered in the environment by several orders of magnitude. If this is true, than much of the broad tuning of receptors as currently characterized by experiment would be functionally inaccessible to the organism. On the other hand, careful experimental work [81, 82] (see Vickers [83] for a review), supported by models of turbulent transport [35, 36], suggests that odorant concentrations fluctuate wildly at the sensory epithelium. Animals may therefore sense odorants quickly during high-concentration bursts, rather than averaging over long timescales. This is supported by behavioral experiments in which flies can execute odor-guided behavior within 100ms [84], as well as the discovery that rodents can identify odors within a similar temporal window [12–14]. Note, however, that these turbulence arguments do not apply to the Drosophila larva, since the fly spends most of its larval stage inside a food source [85].
Even granting that experimental concentrations are unnaturally high, the problem is mitigated by the consideration that all studies necessarily use a limited panel of odorants. As more odorants are tested, more extremely high sensitivities will be filled in for each receptor, and the current picture of combinatorial coding might survive, even if all sensitivities are shifted upwards. Our model, which only considers the relative values of sensitivity and concentration, can say little to resolve this question, which amounts to fixing the mean of the stimulus distribution p(c). But our results do suggest that the variance of p(c) is likely to be high, since this variance is what drives the spread in the optimized sensitivities, and these in turn are a close match to experimental data.
There are several limitations of our receptor-level analysis. First, we have only accounted for excitation of sensory neurons by odorants. This was primarily because inhibition was not detected by calcium imaging in Si et al. [37], which was our primary experimental comparison. In reality, however, antagonistic interactions play a crucial role in mammalian odor coding [38, 86, 87]. Inhibitory responses also emerged in previous theoretical work which optimized over W given a non-zero baseline activity in neurons [24]. Mechanistically, a minimal two step model in which the odorant first binds the receptor with some affinity and then activates it with a different, potentially uncoupled affinity, explained a number of classic observations in psychophysical experiments such as synergy and overshadowing between pairs of odorants [88]. In future work, we hope to optimize over the parameters of this more realistic biophysical model and revisit the above analyses (initial attempts to do so proved numerically unstable within our framework).
Second, lurking in any biophysical model is the question of where to place the nonlinearity. In (1), we have placed the expression matrix E inside of the Hill function. Therefore we are using the Hill function as an empirical fit to ORN activity in the presence of multiple odorants, not as a mechanistic model for coöperativity in odorant receptor signaling. A plausible extension of our model would place E outside of the nonlinearity, then include another nonlinearity for neural activity. To our knowledge, this issue has not been considered in previous works on efficient coding in olfaction, either because they assumed linear processing of the stimulus [25, 41] or canonical olfaction [24, 26] (i.e., a one-hot E whose placement therefore has no effect).
Third, and most fundamentally, biological evolution is a much more constrained and local process than our numerical optimization procedure. For example, real receptors likely cannot change their affinity for one odorant without changing their affinities for others, especially since odorant geometry is a primary determinant of receptor specificity (see Si et al. [37] and del Mármol et al. [58]). Furthermore, organisms develop new receptors from existing ones through a birth-and-death model [7, 69], rather than optimizing a set of randomly initialized receptors as we have done. Any attempt to encode these evolutionary dynamics in our optimization would be computationally challenging, but a principled approach could be fruitful.
With these constraints in mind, it was perhaps surprising that efficient coding could predict any aspects of receptor tuning. Our optimization process becomes more realistic as we compare to more evolutionarily mature receptors that diverged from each other long ago, so the correspondence we found may indicate that the receptors of the larval Drosophila are reasonably mature and tunable.
B. Expression
One key takeaway from our work is that something very close to the one neuron-one receptor rule appears optimal across a broad range of environmental parameters. In any normative model, a chief concern is the extent to which such results depend on specific details of the setup. This is why we have swept over a broad range of parameter values for , and , and inserted multiple models for W. The fairly general emergence of canonical expression in these analyses is likely due in part to the decorrelation of activity (see Figure 3d) that occurs when neurons “spread out” in gene expression space by choosing one receptor. A complete answer, however, would require a more developed theory for the nonlinear setting, which we leave for future work.
Critically, we do not find evidence that receptor coexpression should be finely tuned to either the statistics of receptors or the environment. On the contrary, we find that one neuron-one receptor is basically optimal, and that exceptions to this rule should have varying effects depending on receptor similarity. Stronger claims would need to be weighed against the recent work suggesting that many coexpressed pairs are recent duplicates whose coexpression is driven simply by shared regulatory factors [45, 63]. Ramdya and Benton [89] raise the possibility that such instances may reflect a transient evolutionary state, in which recent duplicates are coexpressed only until regulatory machinery has “caught up” to the duplication event by creating new cell types.
In this respect, our findings differ from a very recent theoretical analysis by Lienkaemper et al. [41], who characterized the dependence of optimal expression on environmental noise and the “signal correlation” using a linear-Gaussian theory. Discrepancies may be due to the nonlinearity we include in the receptor activity or the log-normality of our odor model. It is also important to note that our formulation differs slightly from theirs: whereas they add noise to a Gaussian stimulus and compute , we compute where c is log normal with variance . Thus, any trends that depend on environmental noise cannot be directly compared. On the other hand, in accordance with their theory, we do find that increasing neural noise favors canonical olfaction (see SI Appendix: Fig. S12). Another important difference is that Lienkaemper et al. [41] focus on the setting where the number of neurons L is less than the the number of receptors M, forcing neurons into coexpression if no receptors are to be neglected. In our setting, , which permits neurons to spread out their activity by singly expressing without losing information.
C. Glomeruli
At the glomerular level, the optimized circuit averages across neurons expressing the same receptor, and uses the transfer function to spread out post-synaptic activity. This has already been understood in terms of efficient coding (see Bhandawat et al. [44]), and it is the most intuitive of the three layers. In this way, it serves as a useful sanity check of our computational approach. Because we found canonical expression to be largely robust, we did not attempt to make predictions about glomerular convergence in the case of noncanonical expression.
D. Layer-wise efficient coding
The key structural choice we made in this work was to adopt a layer-wise optimization approach: we first optimized receptor affinities W assuming canonical expression E, then optimized E given fixed W, and finally optimized glomerular connectivity G given fixed E and W. Layer-wise optimization is both computationally convenient, and—as we will argue below—biologically interpretable.
The risk of such a procedure is that the results could depend on the order of the optimizations. As mentioned before, the analysis of canonical expression represents a chicken-or-egg problem. If we optimize W given canonical E, then optimize E given the resulting W, it might be expected that canonical expression would emerge, if this structure is somehow encoded in the optimal W. Such a dependence would severely compromise the generality of our results. We therefore sought to check that the finding of generally-canonical expression was robust to the details of the optimization procedure by optimizing expression for different plausible models of the affinity matrix. For these alternative models, optimization resulted in strictly greater levels of canonical expression than the optimized W. This suggests that the key takeaway—canonical expression is typically optimal—is not an artifact of our optimization procedure.
From a biological perspective, layer-wise optimization may be a reasonable model for the evolution of the convergent motifs we consider. Organisms are not presented with the opportunity to optimize an entire sensory circuit ab initio. Instead, W, E, and G are tuned by different mechanisms, on different timescales, and in different contexts. We therefore had to approximate the constrained setting in which each motif evolved given limited experimental knowledge.
Chemoreceptor families are evolutionarily ancient and continually evolving [2, 6, 7, 30, 56], so we chose to optimize over W first. Since approximate canonical expression is widespread, we plugged in canonical E for this optimization. This also permitted direct comparison to the experimental results presented by Si et al. [37]. Next, it seems reasonable to ask which gene expression programs are optimal given a broad, mature set of receptors. In the mouse, for example, the relative abundance of olfactory receptors can be adapted over the timescale of just a few hours [46, 90], although this process does not generate coordinated non-canonical expression. This framing also addresses the emergence of different expression programs across organisms which each possess mature receptor families. Hence we optimized over E given fixed W.
Glomerular convergence, which constitutes an exquisite example of wiring specificity, emerges within the lifespan of the organism by pruning during development [91, 92]. Since examples of highly adaptive connectivity abound in organisms that learn, we chose to study glomerular convergence in the setting of a mature receptor array and largely canonical expression. Hence we optimized over G given fixed, optimized W and E.
These conceptual distinctions are naturally much cleaner than the underlying biology. For example, receptors continue to evolve after a program of noncanonical expression is established [89]. However, our breakdown enables a tractable model and represents a best guess at the constraints which govern the circuit’s evolution. It is of course possible that layer-wise optimization could yield sensitivities and expression patterns that do not lend themselves to glomerular convergence [93]. However, given access to at least as many neurons as receptor types and canonical expression, the simple intuition that glomerular convergence allows denoising suggests that not much is being lost.
One alternative to our approach is an unconstrained end-to-end optimization of the entire circuit. Wang et al. [94] recovered largely canonical expression and glomerular convergence using unconstrained optimization on a match-to-prototype classification task. This is qualitatively consistent with our results. Their work, however, differs in assuming a particular classification task, and starts with a very different assumption on the stimulus: they directly model W by assuming that the activity of each receptor is independent and uniformly distributed. This differs substantially from the statistics resulting from our optimized sensitivity matrices, and from those measured by Si et al. [37]. This kind of unconstrained optimization can shed light on the computation instantiated by the circuit, but cannot capture the path-dependent quirks of biological evolution [95].
Another alternative would be to formalize our contextspecific model as a joint dynamical system whose dynamics are decomposed across different timescales. Most simply, this would lead to a setting where we optimize W given the expression E that is optimized for each sensitivity matrix, i.e., to study the nested optimization . This however could be too flexible—it seems that many organisms have “committed” to canonical expression, for example, and do not have the ability to smoothly modulate receptor co-expression as receptors evolve. Correspondingly, there may be some biological merit to freezing E while W is optimized.
Fleshing this out completely would require organism-specific knowledge of when and how in evolutionary history each motif emerged. In the absence of such knowledge, we approximate the distinction between the three layers in a way that seems biologically plausible, and treat each one as a separate problem. This piecewise approach limits the normative mathematical claims we can make about the circuit, but it is a closer match to evolutionary dynamics.
E. Dynamics
Our work does not address the dynamics of the odor landscape and its neural representation in the olfactory bulb. As mentioned above, turbulent transport of air to the sensory epithelium induces high-frequency fluctuations in concentration that are richly informative about identity and location [14, 35]. Here, we have ignored these dynamics in order to focus on a statistical problem of compression: namely, the representation of a high-dimensional chemical space by a finite set of receptors. The temporal challenge of encoding rapidly fluctuating odor signals is equally daunting. Recent work has begun to bridge these statistical and dynamical pictures of olfactory sensing [18, 78, 96–101], but many questions remain regarding how to efficiently encode the temporal statistics of the olfactory world. In this vein, it could be interesting to consider minimal extensions of our model that include temporal filtering [51, 102].
F. Conclusion
We reiterate in closing that mutual information is only a rough proxy for performance in olfactory processing. For example, organisms may privilege the sensing of rare but critical odorants, even at the expense of more common ones—this would not be captured by our statistical model [57, 61, 103]. Therefore some of our results are unlikely to transfer perfectly to biological circuits. That said, the great variety of olfactory tasks faced by the typical organism suggests mutual information as a reasonable starting point [23, 24, 103].
The correspondence between our results and experimental measurements suggests that these motifs may constitute uniquely accessible solutions to the problem of processing olfactory stimuli. As chemosensation is further explored in non-model organisms [3, 104], the extent to which canonical principles hold will be an interesting point of focus. While receptors always reflect the organism’s chemical environment, the other layers we consider (receptor expression and the downstream routing of receptor activity) are not obviously bound to any stereotypy. Especially interesting will be cases where departures from canonical olfaction are unambiguous, tunable, and clearly advantageous. Since our simplified modeling suggests that canonical olfaction is optimal, such departures should serve as a flag that something interesting is afoot.
V. METHODS
A. Generating odor mixtures with tunable statistics
We used a Gaussian copula to generate the binary mixture data with tunable mean and covariance. This works as follows.
Draw where .
Set thresholds where is the desired mean vector and ppf is the inverse CDF of the standard normal distribution.
Set .
Given , then has . In general, is not guaranteed to have covariance equal to , but the pairwise covariances are monotonic, nonlinear functions of .
To model a range of frequencies, with most occurring rarely and some occurring frequently, we drew itself from a distribution. We parametrized using a block covariance structure with k blocks, as in Figure 1. Each block represented a mixture in the environment emitting correlated odorants. As we varied k below (down to = 4), it was necessary to manually tune and in order to preserve an approximately fixed odorant frequency distribution while simultaneously achieving the desired correlation structure in the outputs. This is because not all combinations of mean and covariance can be satisfied in binary data.
For low k, we also used a “thinning” step in which, after generating the samples, was set to 0 with probability . This ad hoc intervention was necessary to stabilize the sample-generating algorithm as , ensuring that was approximately constant across different values of k while achieving the desired correlation structure. Thankfully, only minor adjustments to parameters were needed, and exact settings for each k are included as parameter files in the codebase so that data can be generated anew as needed. We include plots of the covariance matrix and frequency distribution for each k in SI Appendix: Fig. S15.
The resulting data represented binarized odorant mixture samples, with typically 10 (range ) odorants per sample. We then assigned an i.i.d. log-normal concentration with variance to each odorant that was present in a given sample. This resulted in an odorant vector which we passed into our model.
To generate samples with a flat frequency distribution (as in SI Appendix: Fig. S10), we repeated the above procedure, but using a constant .
B. Comparison to other models
The stimulus is sensed by the neuron through a panel of m receptors. Each receptor has a vector of affinities for each odorant, which is represented by a receptor-by-odorant sensing matrix W. Our goal is to describe neural activity as a function of odorant concentration. This obviously involves W and c, but what form should it take?
One approach is to assume that each neuron expresses just one receptor, and that there is one neuron per receptor type. We will call this the case of single expression, and denote it . If the firing rate is simply a linear sum of the (Gaussian) stimulus components
where , then can be computed analytically.
However, experimental [37, 38, 86, 88] and theoretical [24, 88] works have shown that saturation and other nonlinear effects play an important role in odor coding, especially when the stimulus varies over orders of magnitude. Thus a more realistic activity model is , where is a Hill function applied element-wise:
| (7) |
We choose a Hill function because experimental work has shown, in the case of monomolecular odorants, that neural activity takes this form as a function of odorant concentration [37], with . See Setup in the main text for more details, and SI Appendix: Fig. S1 for a sweep over Hill coefficients. Note that is sigmoid when c varies over a log scale.
By in (7), we emphasize that this model is making a strong assumption on expression: just one receptor per neuron. Several prior works on efficient coding in olfaction have made this assumption, and within this setup, one can focus on the optimal properties of the sensing matrix W (as in [26], [24]) or how neurons should be allocated to each receptor [25]. This assumption is natural, since many organisms exhibit this organization. But here we wanted to target this for study as well, asking how much of canonical olfaction can be recapitulated by efficient coding theory across both layers W and E. Thus in our model, we include a neuron-by-receptor expression matrix E that allows each neuron to express a variable amount of each receptor. This yields the model (1):
| (8) |
Here is added to model the effect of neural noise. While Poisson noise would more closely model variability in OSN responses [78], differentiating through a Poisson sampler for the purposes of computing gradients was numerically challenging (requiring a continuous relaxation). For simplicity, we therefore used Gaussian noise.
For all analyses except the sweep over neural noise, we set the magnitude of the noise to be . This was to match the approximate mean activity when computed over samples—see Glomeruli and SI Appendix: Fig. S13 for histograms of activity.
The placement of E inside the nonlinearity does not constitute a principled mechanistic model integrating receptor signaling. Instead, we are using the Hill function as an empirical model for neural activity in the presence of mixtures.
C. Details of mutual information maximization
We leverage a variational formulation of mutual information first proposed by Nowozin et al. [55]. Following their presentation, we give a short overview of the theory underlying this approach, which was first presented by Nguyen et al. [105]. The f-divergence between two distributions P and Q defined on a domain X with densities and is defined as
| (9) |
The mutual information is a special case of this: namely the Kullback-Leibler divergence between the joint distribution and the product of marginals , with .
If f is convex and lower semi-continuous, then we can write the Fenchel conjugate
| (10) |
The functions f and are dual in the sense that . Then we can write
| (11) |
| (12) |
| (13) |
where is an appropriately-chosen class of test functions . We have samples from the joint distribution , and we obtain samples from the distribution by shuffling activity. Thus, we can estimate the expectations in (13).
We then simply parameterize using a small neural network, and optimize jointly over the layer in question and the parameters of . Tschannen et al. [106] gave a thorough discussion of critic architectures. Following their recommendations, we used an inner product architecture for so that , where and are separate two-hidden-layer networks with width . To sanity-check that our results did not depend strongly on the choice of critic network, we performed spot-checks using other critic architectures suggested by Tschannen et al. [106].
Such approaches are not without controversy. McAllester and Stratos [52] have argued that, for the Kullback-Leibler divergence (KLD), the expectations in (13) cannot be estimated without exponentially many samples. This is ultimately due to the fact that the KLD has . This theoretical result prompted investigation into why such estimators are still useful in applications such as representation learning [106].
Here, we sidestep this problem by using the Jensen-Shannon divergence (JSD), a symmetrized version of the Kullback-Leibler divergence (KLD) which has , thus avoiding the need for exponentially many samples. The greater stability and robustness of this Jensen-Shannon-based estimator compared to the exact KL formulation were empirically established by Hjelm et al. [53] and Huszár [54]. These variational approaches have gained considerable popularity in the machine learning community, to the point that maximizing mutual information is no longer considered an intractable problem [50, 53, 54].
D. Constrained optimization using mirror descent
When optimizing over W, E, and G, we must account for the constraints on each matrix: W must have positive elements, and both and G must have non-negative elements and have all row sums equal to unity (i.e., each row lies in a simplex). To handle these constraints, we use the framework of mirror descent [107, 108]. For a pedagogical introduction to mirror descent, we direct the reader to Vlad Niculae’s blog post [109].
This approach requires specifying a function which maps from the primal (constrained) space to the dual (unconstrained space), and function which maps from the dual space back to the primal space. So for W, to ensure positivity, we have used . For E and optimizations, we used . Mirror descent is equivalent to using “natural” gradient descent [110] in the dual space, but enjoys favorable numerical properties [108]. It greatly increased the stability of our optimization procedure.
All gradients were computed via automatic differentiation using the JAX library [111]. Initializations were as discussed in the main text: for W, we initialized using a scaled log-normal. For E, we varied the initialization depending on the analysis. In Figure 3, we initialized to random expression. In SI Appendix: Fig. S8, we initialized to canonical expression. Otherwise, we initialized to noncanonical expression. Typically this meant 3-7 receptors per neuron with roughly equal expression, as in Figure 5. However, in SI Appendix: Fig. S11, we initialized to 1-5 receptors per neuron with roughly equal expression. This was to respect the block structure of W when modeling coexpression of correlated receptors. For G, we initialized to random connectivity (Figure 6).
Supplementary Material
ACKNOWLEDGMENTS
We are grateful to A. D. T. Samuel and all the authors of Si et al. [37] for making their data publicly available. J.C.F.C. thanks D. M. Zimmerman and I. Chandok for helpful discussions regarding the evolution and biophysics of receptors. We thank C. Lienkaemper, M. A. Younger, and G. K. Ocker for inspiring discussions regarding non-canonical expression. Moreover, we are indebted to A. D. T. Samuel and the members of his group, and to G. K. Ocker and M. A. Younger, for their thoughtful comments, which have helped us improve our manuscript.
J.C.F.C. was supported by the Harvard Biophysics Graduate Program through training grant T32GM158477 from the National Institutes of Health. F.P. was supported by the Harvard Center for Brain Science (CBS)-NTT Fellowship Program on the Physics of Intelligence. V.N.M. was supported by grant number A47994 from NTT Research, Inc. J.A.Z.-V. was supported by a Junior Fellowship from the Harvard Society of Fellows.
DATA AND CODE AVAILABILITY
Code to perform all optimizations and reproduce all figures is available on GitHub at https://github.com/VNMurthyLab/olfactory-ec.
The dataset of estimated Drosophila larva receptor affinities from Si et al. [37] that we analyzed in Figure 2 was downloaded from A. D. T. Samuel’s lab GitHub, where it is available under an MIT License: https://github.com/samuellab/Larval-ORN/blob/master/Figure3/results/MLEFit.mat.
References
- [1].Hayden S., Bekaert M., Crider T. A., Mariani S., Murphy W. J., and Teeling E. C., Ecological adaptation determines functional mammalian olfactory subgenomes, Genome Research 20, 1–9 (2009). [DOI] [PMC free article] [PubMed] [Google Scholar]
- [2].Valencia-Montoya W. A., Pierce N. E., and Bellono N. W., Evolution of sensory receptors, Annual Review of Cell and Developmental Biology 40, 353–379 (2024). [Google Scholar]
- [3].Yohe L. R. and Brand P., Evolutionary ecology of chemosensation and its role in sensory drive, Current Zoology 64, 525–533 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
- [4].Auer T. O., Khallaf M. A., Silbering A. F., Zappia G., Ellis K., Álvarez-Ocaña R., Arguello J. R., Hansson B. S., Jefferis G. S. X. E., Caron S. J. C., Knaden M., and Benton R., Olfactory receptor and circuit evolution promote host specialization, Nature 579, 402–408 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
- [5].Robertson H. M., Molecular evolution of the major arthropod chemoreceptor gene families, Annual Review of Entomology 64, 227–242 (2019). [Google Scholar]
- [6].Niimura Y. and Nei M., Extensive gains and losses of olfactory receptor genes in mammalian evolution, PLoS ONE 2, e708 (2007). [DOI] [PMC free article] [PubMed] [Google Scholar]
- [7].Nei M., Niimura Y., and Nozawa M., The evolution of animal chemosensory receptor gene repertoires: roles of chance and necessity, Nature Reviews Genetics 9, 951–963 (2008). [Google Scholar]
- [8].Hildebrand J. G. and Shepherd G. M., Mechanisms of olfactory discrimination: Converging evidence for common principles across phyla, Annual Review of Neuroscience 20, 595–631 (1997). [Google Scholar]
- [9].Fulton K. A., Zimmerman D., Samuel A., Vogt K., and Datta S. R., Common principles for odour coding across vertebrates and invertebrates, Nature Reviews Neuroscience 25, 453–472 (2024). [DOI] [PubMed] [Google Scholar]
- [10].Saito H., Chi Q., Zhuang H., Matsunami H., and Mainland J. D., Odor coding by a mammalian receptor repertoire, Science Signaling 2, 10.1126/scisignal.2000016 (2009). [DOI] [Google Scholar]
- [11].Hallem E. A. and Carlson J. R., Coding of odors by a receptor repertoire, Cell 125, 143–160 (2006). [DOI] [PubMed] [Google Scholar]
- [12].Uchida N. and Mainen Z. F., Speed and accuracy of olfactory discrimination in the rat, Nature Neuroscience 6, 1224–1229 (2003). [DOI] [PubMed] [Google Scholar]
- [13].Wilson C. D., Serrano G. O., Koulakov A. A., and Rinberg D., A primacy code for odor identity, Nature Communications 8, 10.1038/s41467-017-01432-4 (2017). [DOI] [Google Scholar]
- [14].Ackels T., Erskine A., Dasgupta D., Marin A. C., Warner T. P. A., Tootoonian S., Fukunaga I., Harris J. J., and Schaefer A. T., Fast odour dynamics are encoded in the olfactory system and guide behaviour, Nature 593, 558 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- [15].Rokni D., Hemmelder V., Kapoor V., and Murthy V. N., An olfactory cocktail party: figure-ground segregation of odorants in rodents, Nature Neuroscience 17, 1225–1232 (2014). [DOI] [PMC free article] [PubMed] [Google Scholar]
- [16].Kundu S., Ganguly A., Chakraborty T. S., Kumar A., and Siddiqi O., Synergism and combinatorial coding for binary odor mixture perception in Drosophila, eneuro 3, ENEURO.0056 (2016). [Google Scholar]
- [17].Riffell J. A., Shlizerman E., Sanders E., Abrell L., Medina B., Hinterwirth A. J., and Kutz J. N., Flower discrimination by pollinators in a dynamic chemical environment, Science 344, 1515–1518 (2014). [DOI] [PubMed] [Google Scholar]
- [18].Jayakumar S., Rigolli N., Mathis M. W., Vergassola M., Mathis A., and Murthy V. N., Mice navigate scent trails using predictive policies, bioRxiv 10.1101/2025.08.27.672631 (2025). [DOI] [Google Scholar]
- [19].Chapuis J. and Wilson D. A., Bidirectional plasticity of cortical pattern recognition and behavioral sensory acuity, Nature Neuroscience 15, 155–161 (2011). [DOI] [PMC free article] [PubMed] [Google Scholar]
- [20].Su C.-Y., Martelli C., Emonet T., and Carlson J. R., Temporal coding of odor mixtures in an olfactory receptor neuron, Proceedings of the National Academy of Sciences 108, 5075–5080 (2011). [Google Scholar]
- [21].Martelli C., Carlson J. R., and Emonet T., Intensity invariant dynamics and odor-specific latencies in olfactory receptor neuron response, The Journal of Neuroscience 33, 6285–6297 (2013). [DOI] [PMC free article] [PubMed] [Google Scholar]
- [22].Linsker R., Self-organization in a perceptual network, Computer 21, 105 (1988). [Google Scholar]
- [23].Simoncelli E. P. and Olshausen B. A., Natural image statistics and neural representation, Annual Review of Neuroscience 24, 1193 (2001). [Google Scholar]
- [24].Qin S., Li Q., Tang C., and Tu Y., Optimal compressed sensing strategies for an array of nonlinear olfactory receptor neurons with and without spontaneous activity, Proceedings of the National Academy of Sciences 116, 20286–20295 (2019). [Google Scholar]
- [25].Tesileanu T., Cocco S., Monasson R., and Balasubramanian V., Adaptation of olfactory receptor abundances for efficient coding, eLife 8, 10.7554/elife.39279 (2019). [DOI] [Google Scholar]
- [26].Zwicker D., Murugan A., and Brenner M. P., Receptor arrays optimized for natural odor statistics, Proceedings of the National Academy of Sciences 113, 5570–5575 (2016). [Google Scholar]
- [27].Krishnamurthy K., Hermundstad A. M., Mora T., Walczak A. M., and Balasubramanian V., Disorder and the neural representation of complex odors, Frontiers in Computational Neuroscience 16, 10.3389/fncom.2022.917786 (2022). [DOI] [Google Scholar]
- [28].Singh V., Tchernookov M., and Balasubramanian V., What the odor is not: Estimation by elimination, Physical Review E 104, 10.1103/physreve.104.024415 (2021). [DOI] [Google Scholar]
- [29].Victor J. D., Boie S. D., Connor E. G., Crimaldi J. P., Ermentrout G. B., and Nagel K. I., Olfactory navigation and the receptor nonlinearity, The Journal of Neuroscience 39, 3713–3727 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
- [30].Benton R., Drosophila olfaction: past, present and future, Proceedings of the Royal Society B: Biological Sciences 289, 10.1098/rspb.2022.2054 (2022). [DOI] [Google Scholar]
- [31].Vosshall L. B., Wong A. M., and Axel R., An olfactory sensory map in the fly brain, Cell 102, 147–159 (2000). [DOI] [PubMed] [Google Scholar]
- [32].Bashkirova E. and Lomvardas S., Olfactory receptor genes make the case for inter-chromosomal interactions, Current Opinion in Genetics I& Development 55, 106–113 (2019). [Google Scholar]
- [33].Brahma A., Frank D. D., Pastor P. D. H., Piekarski P. K., Wang W., Luo J.-D., Carroll T. S., and Kronauer D. J., Transcriptional and post-transcriptional control of odorant receptor choice in ants, Current Biology 33, 5456 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- [34].Zhou Y., O’Connell T. F., Ghaninia M., Smith B. H., Hong E. J., and Sharpee T. O., Low-dimensional olfactory signatures of fruit ripening and fermentation, bioRxiv 10.1101/2024.06.16.599229 (2024). [DOI] [Google Scholar]
- [35].Celani A., Villermaux E., and Vergassola M., Odor landscapes in turbulent environments, Physical Review X 4, 10.1103/physrevx.4.041015 (2014). [DOI] [Google Scholar]
- [36].Nowotny T. and Szyszka P., Dynamics of odor-evoked activity patterns in the olfactory system, in Advances in Dynamics, Patterns, Cognition (Springer International Publishing, 2017) p. 243–261. [Google Scholar]
- [37].Si G., Kanwal J. K., Hu Y., Tabone C. J., Baron J., Berck M., Vignoud G., and Samuel A. D., Structured odorant response patterns across a complete olfactory receptor neuron population, Neuron 101, 950 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
- [38].Zak J. D., Reddy G., Vergassola M., and Murthy V. N., Antagonistic odor interactions in olfactory sensory neurons are widespread in freely breathing mice, Nature Communications 11, 10.1038/s41467-020-17124-5 (2020). [DOI] [Google Scholar]
- [39].Dewan A., Cichy A., Zhang J., Miguel K., Feinstein P., Rinberg D., and Bozza T., Single olfactory receptors set odor detection thresholds, Nature Communications 9, 10.1038/s41467-018-05129-0 (2018). [DOI] [Google Scholar]
- [40].Grosmaitre X., Vassalli A., Mombaerts P., Shepherd G. M., and Ma M., Odorant responses of olfactory sensory neurons expressing the odorant receptor mor23: A patch clamp analysis in gene-targeted mice, Proceedings of the National Academy of Sciences 103, 1970–1975 (2006). [Google Scholar]
- [41].Lienkaemper C., Younger M. A., and Ocker G. K., When non-canonical olfaction is optimal, Biorxiv 10.1101/2025.02.27.640624 (2025). [DOI] [Google Scholar]
- [42].Potter S. M., Zheng C., Koos D. S., Feinstein P., Fraser S. E., and Mombaerts P., Structure and emergence of specific olfactory glomeruli in the mouse, The Journal of Neuroscience 21, 9713–9723 (2001). [DOI] [PMC free article] [PubMed] [Google Scholar]
- [43].Herre M., Goldman O. V., Lu T.-C., Caballero-Vidal G., Qi Y., Gilbert Z. N., Gong Z., Morita T., Rahiel S., Ghaninia M., Ignell R., Matthews B. J., Li H., Vosshall L. B., and Younger M. A., Non-canonical odor coding in the mosquito, Cell 185, 3104 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- [44].Bhandawat V., Olsen S. R., Gouwens N. W., Schlief M. L., and Wilson R. I., Sensory processing in the Drosophila antennal lobe increases reliability and separability of ensemble odor representations, Nature Neuroscience 10, 1474–1482 (2007). [DOI] [PMC free article] [PubMed] [Google Scholar]
- [45].Adavi E. D., dos Anjos V. L., Kotb S., Metz H. C., Tian D., Zhao Z., Zung J. L., Rose N. H., and McBride C. S., Olfactory receptor coexpression and co-option in the dengue mosquito, bioRxiv 10.1101/2024.08.21.608847 (2024). [DOI] [Google Scholar]
- [46].Tsukahara T., Brann D. H., Pashkovski S. L., Guitchounts G., Bozza T., and Datta S. R., A transcriptional rheostat couples past activity to future sensory responses, Cell 184, 6326 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- [47].Bintu B., Isogai Y., Zhuang X., and Dulac C., Molecular and spatial organization of the primary olfactory system and its responses to social odors, bioRxiv 10.1101/2025.05.02.651832 (2025). [DOI] [Google Scholar]
- [48].Brann D. H., Tsukahara T., Tau C., Kalloor D., Lubash R., Thamarai Kannan L., Klimpert N., Kollo M., Escamilla-Del-Arenal M., Bintu B., Bozza T., and Datta S. R., A spatial code governs olfactory receptor choice and aligns sensory maps in the nose and brain, bioRxiv 10.1101/2025.05.02.651738 (2025). [DOI] [Google Scholar]
- [49].Shao S., Meister M., and Gjorgjieva J., Efficient population coding of sensory stimuli, Phys. Rev. Res. 5, 043205 (2023). [Google Scholar]
- [50].Abdelaleem E., Martini K. M., and Nemenman I., Accurate estimation of mutual information in high dimensional data, arXiv 10.48550/ARXIV.2506.00330 (2025). [DOI] [Google Scholar]
- [51].Jun N. Y., Field G., and Pearson J., Efficient coding, channel capacity, and the emergence of retinal mosaics, in Advances in Neural Information Processing Systems, Vol. 35, edited by Koyejo S., Mohamed S., Agarwal A., Belgrave D., Cho K., and Oh A. (Curran Associates, Inc., 2022) pp. 32311–32324. [PMC free article] [PubMed] [Google Scholar]
- [52].McAllester D. and Stratos K., Formal limitations on the measurement of mutual information, in Proceedings of the Twenty Third International Conference on Artificial Intelligence and Statistics, Proceedings of Machine Learning Research, Vol. 108, edited by Chiappa S. and Calandra R. (PMLR, 2020) pp. 875–884. [Google Scholar]
- [53].Hjelm R. D., Fedorov A., Lavoie-Marchildon S., Grewal K., Bachman P., Trischler A., and Bengio Y., Learning deep representations by mutual information estimation and maximization, in International Conference on Learning Representations (2019). [Google Scholar]
- [54].Huszár F., How (not) to train your generative model: Scheduled sampling, likelihood, adversary?, arXiv 10.48550/ARXIV.1511.05101 (2015). [DOI] [Google Scholar]
- [55].Nowozin S., Cseke B., and Tomioka R., f-gan: Training generative neural samplers using variational divergence minimization, in Advances in Neural Information Processing Systems, Vol. 29, edited by Lee D., Sugiyama M., Luxburg U., Guyon I., and Garnett R. (Curran Associates, Inc., 2016). [Google Scholar]
- [56].Bear D. M., Lassance J.-M., Hoekstra H. E., and Datta S. R., The evolving neural and genetic architecture of vertebrate olfaction, Current Biology 26, R1039–R1049 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
- [57].Sakurai T., Mitsuno H., Mikami A., Uchino K., Tabuchi M., Zhang F., Sezutsu H., and Kanzaki R., Targeted disruption of a single sex pheromone receptor gene completely abolishes in vivo pheromone response in the silkmoth, Scientific Reports 5, 10.1038/srep11001 (2015). [DOI] [Google Scholar]
- [58].del Mármol J., Yedlin M. A., and Ruta V., The structural basis of odorant recognition in insect olfactory receptors, Nature 597, 126–131 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- [59].Meyerhof W., Batram C., Kuhn C., Brockhoff A., Chudoba E., Bufe B., Appendino G., and Behrens M., The molecular receptive ranges of human tas2r bitter taste receptors, Chemical Senses 35, 157–170 (2009). [DOI] [PubMed] [Google Scholar]
- [60].Pregitzer P., Schubert M., Breer H., Hansson B. S., Sachse S., and Krieger J., Plant odorants interfere with detection of sex pheromone signals by male heliothis virescens, Frontiers in Cellular Neuroscience 6, 10.3389/fncel.2012.00042 (2012). [DOI] [Google Scholar]
- [61].Renou M., Party V., Rouyar A., and Anton S., Olfactory signal coding in an odor background, Biosystems 136, 35–45 (2015). [DOI] [PubMed] [Google Scholar]
- [62].Namiki S., Iwabuchi S., and Kanzaki R., Representation of a mixture of pheromone and host plant odor by antennal lobe pro jection neurons of the silkmoth bombyx mori, Journal of Comparative Physiology A 194, 501–515 (2008). [Google Scholar]
- [63].Mika K. and Benton R., Olfactory receptor gene regulation in insects: Multiple mechanisms for singular expression, Frontiers in Neuroscience 15, 10.3389/fnins.2021.738088 (2021). [DOI] [Google Scholar]
- [64].Glotzer G. L., Pastor P. D. H., and Kronauer D. J. C., Transcriptional interference gates monogenic odorant receptor expression in ants, bioRxiv 10.1101/2025.08.21.671318 (2025). [DOI] [Google Scholar]
- [65].Olshausen B. A. and Field D. J., Emergence of simple-cell receptive field properties by learning a sparse code for natural images, Nature 381, 607–609 (1996). [DOI] [PubMed] [Google Scholar]
- [66].Task D., Lin C.-C., Vulpe A., Afify A., Ballou S., Brbic M., Schlegel P., Raji J., Jefferis G. S., Li H., Menuz K., and Potter C. J., Chemoreceptor co-expression in Drosophila melanogaster olfactory neurons, eLife 11, 10.7554/elife.72599 (2022). [DOI] [Google Scholar]
- [67].Goldman A. L., Van der Goes van Naters W., Lessing D., Warr C. G., and Carlson J. R., Coexpression of two functional odor receptors in one neuron, Neuron 45, 661–666 (2005). [DOI] [PubMed] [Google Scholar]
- [68].Lebreton S., Borrero-Echeverry F., Gonzalez F., Solum M., Wallin E. A., Hedenström E., Hansson B. S., Gustavsson A.-L., Bengtsson M., Birgersson G., Walker W. B., Dweck H. K. M., Becher P. G., and Witzgall P., A drosophila female pheromone elicits species-specific long-range attraction via an olfactory channel with dual specificity for sex and food, BMC Biology 15, 10.1186/s12915-017-0427-x (2017). [DOI] [Google Scholar]
- [69].Sánchez-Gracia A., Vieira F. G., and Rozas J., Molecular evolution of the major chemosensory gene families in insects, Heredity 103, 208–216 (2009). [DOI] [PubMed] [Google Scholar]
- [70].Adipietro K. A., Mainland J. D., and Matsunami H., Functional evolution of mammalian odorant receptors, PLoS Genetics 8, e1002821 (2012). [DOI] [PMC free article] [PubMed] [Google Scholar]
- [71].Malnic B., Godfrey P. A., and Buck L. B., The human olfactory receptor gene family, Proceedings of the National Academy of Sciences 101, 2584–2589 (2004). [Google Scholar]
- [72].Yang J.-Y., O’Connell T. F., Hsu W.-M. M., Bauer M. S., Dylla K. V., Sharpee T. O., and Hong E. J., Restructuring of olfactory representations in the fly brain around odor relationships in natural sources, bioRxiv 10.1101/2023.02.15.528627 (2023). [DOI] [Google Scholar]
- [73].Mori K., Nagao H., and Yoshihara Y., The olfactory bulb: Coding and processing of odor molecule information, Science 286, 711–715 (1999). [DOI] [PubMed] [Google Scholar]
- [74].Mombaerts P., Molecular biology of odorant receptors in vertebrates, Annual Review of Neuroscience 22, 487–509 (1999). [Google Scholar]
- [75].Laughlin S., A simple coding procedure enhances a neuron’s information capacity, Zeitschrift für Naturforschung C 36, 910 (1981). [Google Scholar]
- [76].Baraniuk R., Compressive sensing [lecture notes], IEEE Signal Processing Magazine 24, 118–121 (2007). [Google Scholar]
- [77].Blumensath T., Compressed sensing with nonlinear observations and related nonlinear optimization problems, IEEE Transactions on Information Theory 59, 3466 (2013). [Google Scholar]
- [78].Zavatone-Veth J., Masset P., Tong W., Zak J. D., Murthy V., and Pehlevan C., Neural circuits for fast Poisson compressed sensing in the olfactory bulb, in Advances in Neural Information Processing Systems , Vol. 36, edited by Oh A., Naumann T., Globerson A., Saenko K., Hardt M., and Levine S. (Curran Associates, Inc., 2023) pp. 64793–64828. [PMC free article] [PubMed] [Google Scholar]
- [79].Malnic B., Hirono J., Sato T., and Buck L. B., Combinatorial receptor codes for odors, Cell 96, 713–723 (1999). [DOI] [PubMed] [Google Scholar]
- [80].Wachowiak M., Dewan A., Bozza T., O’Connell T. F., and Hong E. J., Recalibrating olfactory neuroscience to the range of naturally occurring odor concentrations, Journal of Neuroscience 45, 10.1523/JNEUROSCI.1872-24.2024 (2025). [DOI] [Google Scholar]
- [81].Murlis J. and Jones C. D., Fine-scale structure of odour plumes in relation to insect orientation to distant pheromone and other attractant sources, Physiological Entomology 6, 71–86 (1981). [Google Scholar]
- [82].Mylne K. R. and Mason P. J., Concentration fluctuation measurements in a dispersing plume at a range of up to 1000 m, Quarterly Journal of the Royal Meteorological Society 117, 177–206 (1991). [Google Scholar]
- [83].Vickers N. J., Mechanisms of animal navigation in odor plumes, The Biological Bulletin 198, 203 (2000). [DOI] [PubMed] [Google Scholar]
- [84].van Breugel F. and Dickinson M., Plume-tracking behavior of flying Drosophila emerges from a set of distinct sensory-motor reflexes, Current Biology 24, 274–286 (2014). [DOI] [PubMed] [Google Scholar]
- [85].Kreher S. A., Kwon J. Y., and Carlson J. R., The molecular basis of odor coding in the Drosophila larva, Neuron 46, 445–456 (2005). [DOI] [PubMed] [Google Scholar]
- [86].Pfister P., Smith B. C., Evans B. J., Brann J. H., Trimmer C., Sheikh M., Arroyave R., Reddy G., Jeong H.-Y., Raps D. A., Peterlin Z., Vergassola M., and Rogers M. E., Odorant receptor inhibition is fundamental to odor encoding, Current Biology 30, 2574 (2020). [DOI] [PubMed] [Google Scholar]
- [87].Xu L., Li W., Voleti V., Zou D.-J., Hillman E. M. C., and Firestein S., Widespread receptor-driven modulation in peripheral olfactory coding, Science 368, 10.1126/science.aaz5390 (2020). [DOI] [Google Scholar]
- [88].Reddy G., Zak J. D., Vergassola M., and Murthy V. N., Antagonism in olfactory receptor neurons and its implications for the perception of odor mixtures, eLife 7, 10.7554/elife.34958 (2018). [DOI] [Google Scholar]
- [89].Ramdya P. and Benton R., Evolving olfactory systems on the fly, Trends in Genetics 26, 307–316 (2010). [DOI] [PubMed] [Google Scholar]
- [90].Santoro S. W. and Dulac C., The activity-dependent histone variant h2be modulates the life span of olfactory neurons, eLife 1, e00070 (2012). [DOI] [PMC free article] [PubMed] [Google Scholar]
- [91].Hong W. and Luo L., Genetic control of wiring specificity in the fly olfactory system, Genetics 196, 17–29 (2014). [DOI] [PMC free article] [PubMed] [Google Scholar]
- [92].Zou D.-J., Feinstein P., Rivers A. L., Mathews G. A., Kim A., Greer C. A., Mombaerts P., and Firestein S., Postnatal refinement of peripheral olfactory projections, Science 304, 1976–1979 (2004). [DOI] [PubMed] [Google Scholar]
- [93].Gutierrez G. J., Rieke F., and Shea-Brown E. T., Nonlinear convergence boosts information coding in circuits with parallel outputs, Proceedings of the National Academy of Sciences 118, 10.1073/pnas.1921882118 (2021). [DOI] [Google Scholar]
- [94].Wang P. Y., Sun Y., Axel R., Abbott L., and Yang G. R., Evolving the olfactory system with machine learning, Neuron 109, 3879 (2021). [DOI] [PubMed] [Google Scholar]
- [95].Cisek P. and Hayden B. Y., Neuroscience needs evolution, Philosophical Transactions of the Royal Society B: Biological Sciences 377, 20200518 (2022). [Google Scholar]
- [96].Kadakia N. and Emonet T., Front-end Weber-Fechner gain control enhances the fidelity of combinatorial odor coding, eLife 8, e45293 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
- [97].Choi K., Rosenbluth W., Graf I. R., Kadakia N., and Emonet T., Bifurcation enhances temporal information encoding in the olfactory periphery, PRX Life 2, 10.1103/prxlife.2.043011 (2024). [DOI] [Google Scholar]
- [98].Giaffar H., Shuvaev S., Rinberg D., and Koulakov A. A., The primacy model and the structure of olfactory space, PLOS Computational Biology 20, 1 (2024). [Google Scholar]
- [99].Bourassa F. X. P., Francois P., Reddy G., and Vergassola M., Manifold learning for olfactory habituation to strongly fluctuating backgrounds, bioRxiv 10.1101/2025.05.26.656161 (2025). [DOI] [Google Scholar]
- [100].Chapochnikov N. M., Pehlevan C., and Chklovskii D. B., Normative and mechanistic model of an adaptive circuit for efficient encoding and feature extraction, Proceedings of the National Academy of Sciences 120, e2117484120 (2023). [Google Scholar]
- [101].Reddy G., Shraiman B. I., and Vergassola M., Sector search strategies for odor trail tracking, Proceedings of the National Academy of Sciences 119, e2107431118 (2022). [Google Scholar]
- [102].Golkar S., Berman J., Lipshutz D., Haret R. M., Gollisch T., and Chklovskii D. B., Neuronal temporal filters as normal mode extractors, Physical Review Research 6, 013111 (2024). [Google Scholar]
- [103].Durian S. C. L., Bojanek K., Marre O., and Palmer S. E., Preserving predictive information under biologically plausible compression, bioRxiv 10.1101/2025.03.12.642864 (2025). [DOI] [Google Scholar]
- [104].van Giesen L., Kilian P. B., Allard C. A. H., and Bellono N. W., Molecular basis of chemotactile sensation in octopus, Cell 183, 594 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
- [105].Nguyen X., Wainwright M. J., and Jordan M., Estimating divergence functionals and the likelihood ratio by penalized convex risk minimization, in Advances in Neural Information Processing Systems, Vol. 20, edited by Platt J., Koller D., Singer Y., and Roweis S. (Curran Associates, Inc., 2007). [Google Scholar]
- [106].Tschannen M., Djolonga J., Rubenstein P. K., Gelly S., and Lucic M., On mutual information maximization for representation learning, in International Conference on Learning Representations (2020). [Google Scholar]
- [107].Beck A. and Teboulle M., Mirror descent and nonlinear projected subgradient methods for convex optimization, Operations Research Letters 31, 167 (2003). [Google Scholar]
- [108].Raskutti G. and Mukherjee S., The information geometry of mirror descent, IEEE Transactions on Information Theory 61, 1451 (2015). [Google Scholar]
- [109].Niculae V., Optimizing with constraints: reparametrization and geometry, online blog post (2020), URL: https://vene.ro/blog/mirror-descent.html. [Google Scholar]
- [110].Amari S.-i., Natural gradient works efficiently in learning, Neural Computation 10, 251 (1998). [Google Scholar]
- [111].Bradbury J., Frostig R., Hawkins P., Johnson M. J., Leary C., Maclaurin D., Necula G., Paszke A., Van-derPlas J., Wanderman-Milne S., and Zhang Q., JAX: composable transformations of Python+NumPy programs (2018). [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Data Availability Statement
Code to perform all optimizations and reproduce all figures is available on GitHub at https://github.com/VNMurthyLab/olfactory-ec.
The dataset of estimated Drosophila larva receptor affinities from Si et al. [37] that we analyzed in Figure 2 was downloaded from A. D. T. Samuel’s lab GitHub, where it is available under an MIT License: https://github.com/samuellab/Larval-ORN/blob/master/Figure3/results/MLEFit.mat.






