Skip to main content
NIHPA Author Manuscripts logoLink to NIHPA Author Manuscripts
. Author manuscript; available in PMC: 2026 Aug 23.
Published in final edited form as: Nat Neurosci. 2026 Apr 16;29(5):1068–1078. doi: 10.1038/s41593-026-02256-6

Spatial, temporal, and Notch determination of terminal selector expression controls neuronal cell fate in the Drosophila optic lobe

Félix Simon 1,2,*, Isabel Holguera 1,2, Yen-Chung Chen 1, Jennifer Malin 1, Priscilla Valentino 3,4, Claire Njoo-Deplante 1, Rana Naja El-Danaf 5, Katarina Kapuralin 5,6, Ted Erclik 3,4, Nikolaos Konstantinides 2, Mehmet Neset Özel 7, Claude Desplan 1,5,*
PMCID: PMC13499599  NIHMSID: NIHMS2196385  PMID: 41992006

Abstract

In the medulla of the Drosophila optic lobe, the identity of each neuronal type is specified in progenitors and newborn neurons via the integration of temporal, spatial, and Notch-driven patterning mechanisms. This identity is maintained in differentiating and adult neurons by the continuous expression of neuronal type-specific combinations of transcription factors called terminal selectors, which are thought to control all neuronal type-specific features. How the patterning mechanisms establish terminal selector expression is unknown. Here, we have used single-cell mRNA-sequencing to characterize the spatial origins of medulla neurons. Combined to our previous characterization of their temporal and Notch origins, this allowed us to identify correlations between patterning information, terminal selector expression and neuronal features. Our results suggest that different subsets of the patterning information accessible to a given neuronal type control the expression of each of its terminal selectors and of modules of terminal features, including neurotransmitter identity.

Introduction

Brain functions rely on communication between specialized neuronal types, each characterized by molecular and morphological features. Understanding how these features are established is key to understanding brain development. We explored this question in the Drosophila optic lobe, which is constituted of four neuropils (Fig. 1a): the lamina, medulla, lobula and lobula plate. Each is subdivided into ~800 columns that correspond to the ~800 ommatidia (i.e., unit eyes) of the retina, and in several layers orthogonal to these columns, except for the lamina that has no layers. The stereotyped development and repetitive structure of the Drosophila optic lobe has enabled the comprehensive characterization of its neuronal morphological diversity1–4 and of its synaptic connectivity5–10. We and others have also characterized its transcriptomic diversity11–15: we performed single-cell mRNA sequencing (scRNA-seq) at six different developmental stages11, identified ~170 neuronal clusters, and matched almost 90 of these to known morphological neuronal types11,16–21. Lastly, the factors responsible for the specification of optic lobe neuronal types have been characterized in great detail18,20,22,23.

Fig. 1: Specification of neuronal diversity in the Drosophila optic lobe.

Fig. 1:

a) Cross-section of the adult visual system, with neuropil labels located adjacent to each one. Optic lobe neuron arbors can be restricted to a unique column (narrow-field: the two neuronal types on the left, including a photoreceptor), or they can span multiple columns (wide-field, right). Dashed lines: boundaries between neuropil layers.

b) Schematic of the third-instar larval optic lobe, containing the lateral OPC (green), main OPC (blue), tips of the OPC (tOPC, orange) and the inner proliferation center (IPC, purple).

c) Lateral view of the main OPC, showing the domain of expression of the spatial specification factors: Vsx (dark blue, anterior), Optix (light blue, middle), Dpp (green, posterior). The tips of the OPC (tOPC) express Wg (orange). Disco and Hh (pink hatched line) define the ventral domain while Spalt (yellow hatched lines) defines the dorsal domain. NB: neuroblast.

a-c) A: anterior, D: dorsal, L: lateral, M: medial, P: posterior, V: ventral.

d) In the main OPC, a series of temporal TFs (blue) defines temporal windows (grey columns) in neuroblasts, leading to the expression of “concentric genes” (green) in neurons (adapted from Konstantinides et al.28).

e) Diagram of Notch signaling in optic lobe neurons. GMC: ganglion mother cell.

Optic lobe neurons are produced from two crescent-shaped neuroepithelia at the third-instar larval stage (L3, Fig. 1b), the inner proliferation center and the outer proliferation center (OPC). The OPC can be further subdivided into the lateral OPC that produces lamina neurons, the tips of the OPC that express wingless (wg), and the main OPC that produces medulla neurons (Fig. 1b). We focused on the main OPC, which generates the greatest neuronal diversity and for which neuronal specification is best understood. At the beginning of the L3 stage, a neurogenic wave starts converting the OPC neuroepithelial cells into neural stem cells called neuroblasts in Drosophila. Neuroblasts divide multiple times to self-renew and produce an intermediate progenitor called a ganglion mother cell, which in turn divides once more to generate two neurons. Neuronal identity is specified by the integration of three mechanisms24–29. First (Fig. 1c), the main OPC neuroepithelium is subdivided into different spatial domains that express either the spatial transcription factors (TFs) Visual system homeobox 1 (Vsx1) or Optix, or the signaling molecule Decapentaplegic (Dpp)25. Additionally, the dorsal part of the main OPC neuroepithelium expresses Spalt TFs (Salm and Salr), while the ventral side expresses Disco TFs (Disco and Disco-r) and, only early in development, Hedgehog (Hh)24,30. We will hereafter refer to ventral and dorsal Vsx, Optix, and Dpp domains as v/dVsx, v/dOptix and v/dDpp, respectively. Second (Fig. 1d), all main OPC neuroblasts sequentially express, as they age, a cascade of temporal transcription factors (tTFs) whose expression partially overlaps25–29. Although tTFs often do not maintain their expression in neurons, several downstream TFs termed “concentric genes” (Fig. 1d) are expressed in neuronal types emerging from one or several juxtaposed temporal windows, and are therefore expressed in a concentric fashion in the main OPC28,31 (reflecting the spatial organization of neuroblasts of identical age in the main OPC, Fig. 1c). Third (Fig. 1e), the Notch pathway is transiently active in one of the two neurons produced by each ganglion mother cell (NotchON neurons) but inactive in the other (NotchOFF neurons)11,26. We hereafter use “specification factor” to refer to any factor involved in spatial, temporal or Notch patterning, and “developmental origin” for any combination of spatial, temporal and Notch origin. Moreover, because neurons are produced in temporal windows that include overlapping expression of several tTFs, and from spatial origins (e.g., dVsx and dOptix) that can encompass the domain of several spatial specification factors (e.g., Vsx, Optix, Salm), we hereafter use “specification module” as a general term that refers to any temporal window, spatial origin, or Notch status that produces a neuronal type.

Spatial and temporal specification factors are mostly not maintained in neurons, and the direct effectors of Notch signaling are only present in newborn neurons11,26. It has been proposed that the information encoded by the specification factors is propagated to differentiating neurons by downstream TFs called terminal selectors (TSs), which are expressed in neuronal type-specific combinations maintained during development and adulthood, and which control all specific features of a neuronal type by binding the cis-regulatory regions of either terminal effector genes or other TFs controlling terminal effector genes32,33. Patterning mechanisms establishing TS expression were identified for instance for ttx-334–36 in C. elegans and Pet-137,38 in mice, but high-throughput studies of the establishment of TS expression are lacking. We previously identified TS candidates (TSCs) in the optic lobe: combinations of TFs with an average of 10 per neuron that are unique to each neuronal type and are maintained throughout neuronal life11,17. Although in most cases the control of terminal effector genes by our TSCs was not demonstrated, this set of candidates should contain all genuine TSs even if some candidates do not act as bona fide TSs. Notably, we were able to convert distinct neuronal types that shared very similar TSC combinations into one another by rendering these combinations identical17.

Although we recently described the temporal origin and the Notch status of all the main OPC clusters in our scRNA-seq atlas28, the lack of knowledge of their spatial origin has prevented us from exploring how the specification factors establish together the TS combinations specific to each neuronal type. In this study, we used a scRNA-seq approach to comprehensively characterize the spatial origin of main OPC neurons. We then identified how each temporal window, spatial domain, Notch status, or any combination of these three specification mechanisms, correlates with TSC expression (as well as other genes specifically expressed in neuronal types and morphological features) in differentiating and adult neuronal types. Lastly, we trained random forest models that show that the developmental origin of optic lobe neurons is highly predictive of their TSC expression, and that developmental origin or TSC expression are equally informative in predicting gene expression or morphological features of optic lobe neurons.

Results

Profiling the neuronal types produced in each spatial domain

Identifying the spatial origin of neuronal types has traditionally relied on low-throughput approaches20,25,39 (immunostainings, lineage tracing, etc.) that are often difficult to implement due to the lack of appropriate marker genes, antibodies, or driver lines. To identify the spatial origins of all main OPC neurons, we drove nuclear GFP expression in cells produced from various main OPC spatial domains using domain-specific Gal4 drivers, isolated the labeled cells by Fluorescence-Activated Cell Sorting (FACS), performed scRNA-seq, and annotated the cells from each population using a neural network classifier trained on our previously published optic lobe scRNA-seq atlas11 (Fig. 2a). When it was technically possible, we sequenced adult neurons by using memory cassettes that label all progeny of cells that expressed a given Gal4 driver earlier in development (Fig. 2b, Extended Data Fig. 1a, b), because at early pupal stages some neurons are still too immature to be identified precisely by our classifier11. We tested the specificity of the lines by checking their GFP expression pattern before the rearrangement of neuronal somas31, i.e., either at the L3 larval stage or between 0 to 12.5% of pupal development (P0-P12.5, Fig. 2b, c). All lines labelled the expected main OPC populations and sometimes some cells inside or outside the main OPC, which we accounted for in later analyses (Extended Data Fig. 1c–e, Methods). We used pxb-Gal4 rather than Vsx-Gal427 because Vsx is also expressed later in neurons from the v/dOptix domains (Extended Data Fig. 1c, f).

Fig. 2: Profiling the neuronal types produced in each spatial domain.

Fig. 2:

a) Schematic of the approach used to identify the neurons produced in each main OPC domain. Gal4 lines driving the expression of a nuclear GFP were used to isolate cells from a given spatial origin with FACS, either in early pupae (top) when the Gal4 is expressed in neuronal progenitors, or in adults (bottom) using a memory cassette (MC) to label permanently the neurons produced in a given spatial domain (Extended Data Fig. 1a). Sorted cells were then sequenced, and the single-cell transcriptomes were annotated using a neural network classifier11.

b) Summary of the stages at which we tested the expression of, or sequenced, each of the lines used. The expression pattern of the lines used was tested before migration of cell bodies, either at the L3 stage or in early pupa. Temperature sensitive memory cassettes (ts-MC) are inactive at 18°C, but permanently label the cells expressing them at 29°C as well as their progeny.

c) Maximum intensity projections showing that each of the lines sequenced labelled the expected main OPC populations (see Fig. 1c), following the protocol summarized on panel B. Labelled neurons outside of the dashed lines are not part of the main OPC. Schematics show the orientation of the maximal intensity projection.

After performing scRNA-seq at the stages indicated on Fig. 2b, we used our classifier11 to annotate the transcriptome of each cell with a confidence score. We did not discard the “Classes” (defined hereafter as a group of cells with the same annotation) annotated with a lower confidence because this could be due to biological reasons (Extended Data Fig. 2a), but flagged them for later analyzes (Supplementary Data 1, Methods). We then normalized neuronal abundances to be able to compare them between different datasets (Extended Data Fig. 2b).

Although the spatial origin was very clear for most neuronal types (Extended Data Fig. 2c, Supplementary Data 1), we sometimes found a low number of cells belonging to a neuronal type from a known spatial origin in datasets where they were predicted to be absent. This can be explained by experimental (e.g. imperfect sorting), bioinformatic (e.g. imperfect filtering), or biological (e.g. cell types with very similar transcriptomes, which are difficult to distinguish using the neural network) variation.

Assigning a spatial origin to each optic lobe neuronal type

To validate our scRNA-seq results we used the known spatial origins of Dm839, Mi126,31, Pm1/2/325, T140 and TEv/d11 neurons, and experimentally determined the spatial origins of 32 additional clusters. To do so, we identified markers expressed early in these clusters (Extended Data Fig. 3a) and assessed the location of their expression by immunostaining in the main OPC cortex at late L3 stage (before cell bodies rearrange) (Fig. 3a, b, Extended Data Fig. 3b–j, Supplementary Table 1). We then computed the enrichment of each Class in each dataset (Extended Data Fig. 2b) by dividing the normalized abundance of each Class by the normalized abundance in our previously published scRNA-seq atlas, and binarized these enrichment values into “presence” or “absence” of a Class in a dataset by setting a threshold based on experimentally validated spatial origins (Extended Data Fig. 4, 5, Supplementary Data 1). After binarization, the spatial origins determined by scRNA-seq matched the ones determined experimentally for 39 out of 40 neuronal types (Extended Data Fig. 4). They did not completely match for Dm11, but this could be due to mis-annotations of transcriptionally similar cell types by the neural network (see Extended Data Fig. 6 and Supplementary Table 1). Moreover, our results suggest that either Pm1 (cluster 163) is produced in vOptix, and not vDpp as previously thought25, or that Pm1 corresponds to cluster 180 (Supplementary Table 1). Lastly, we generated a table assigning a spatial origin to all other main OPC clusters by considering normalized abundances, proximity to enrichment binarization thresholds, UMAP representation, annotation confidence and experimental results. This allowed us to assign a spatial origin for 93 neuronal clusters, including 10 Dm neuronal clusters that we discussed previously20, in addition to the 8 previously known. Each assignment is explained in Supplementary Table 1, while Fig. 3c represents a summarized view of this table.

Fig. 3: Identification of the spatial origins of main OPC neurons.

Fig. 3:

a, b, d) Lateral view of the main OPC at the L3 stage. Grey dashed line: outline of the optic lobe. Scale bars are 20 μm.

a) Mi9 neurons (white dashed line) are Svp+ D+ (Extended Data Fig. 3a) and localized in the v/dOptix and v/dDpp domains.

b) Toy+ Tup+ cells, which include Dm10 neurons (Extended Data Fig. 3a), are localized in the vVsx (i), vOptix (i’), and vDpp (i’’) domains (arrows).

c) Developmental origin (spatial, temporal, and Notch) and abundance of main OPC neuronal clusters. Clusters are indicated as “unconfirmed” when they were found in a dataset containing both dorsal and ventral cells (ex: dpp dataset), and in a dataset containing only ventral cells (ex: hh dataset), but we lack a dorsal dataset to confirm their presence in the corresponding dorsal domain. Uncertainties in spatial origin assignments are explained in Supplementary Table 1. The names of annotated clusters are represented in bold. TO: temporal origin, NO: Notch origin, Size: number of cells comprised in the corresponding cluster in our adult scRNA-seq atlas11.

d) Tm9 neurons (white dashed line) are the first row of Dac+ Vvl+ neurons (Extended Data Fig. 3a). Tm9v but not Tm9d is produced from the Vsx domain, in the vVsx stripe.

e) Additional spatial domains of the main OPC, lateral view.

Our scRNA-seq data suggested the existence of two additional spatial domain subdivisions. First, according to our sequencing results, Tm9d was produced from the dOptix and dDpp domains and Tm9v from the vVsx and vOptix domains (Supplementary Table 1). Such asymmetry for the dorsal and ventral subtypes of the same neuronal type was unexpected, but we experimentally validated that Tm9v is produced in a small subregion of the vVsx domain at the border with the vOptix domain (Fig. 3d, e, hereafter called “vVsx stripe”) while Tm9d is absent from the Vsx domain. Second, although Dm1 and Dm12 shared markers expressed in the vDpp domain (Extended Data Fig. 3g), only Dm12 was also found in the dpp dataset (Supplementary Data 1). We experimentally validated that there exists a domain that produces Dm12 but not Dm1, where Optix and dpp overlap (Dpp/Optix domain), and showed that this domain is partly established by dpp signaling20. This experimental validation of two subdomains predicted from our scRNA-seq data (Fig. 3e) further emphasizes the accuracy of our results. However, it also means that neurons present in both the Optix dataset and the dpp dataset must be produced from the Dpp+ Optix+ domain, but whether they are also produced from the Dpp+ Optix− domain is unclear. For instance, we assigned Mi1 and Tm1/4/6 as being produced in the whole main OPC, but some evidence (Supplementary Table 1) suggests that they may not be produced in the Dpp+ Optix− domain.

Moreover, medulla neurons can be present in one copy per column, and are called synperiodic, or less than once per column and are called infraperiodic. We previously showed that the NotchON progeny of the Hth temporal window produces the synperiodic neuron Mi1 from the whole main OPC, while the NotchOFF progeny produces several infraperiodic neuronal types from restricted spatial domains18,25, and we suggested that NotchON neurons might ignore spatial patterning while NotchOFF neurons would respond to spatial information. Consistently, we have previously shown that spatial domain size correlates with the abundance of Dm neurons (which are NotchOFF)20, and our results show that in early temporal windows NotchON neurons are indeed synperiodic and produced from the whole main OPC. However, in late temporal windows, several synperiodic neurons are either NotchOFF or produced from restricted spatial domains. This was already known for Dm8, an abundant neuron produced from a small spatial domain, and we previously suggested that Dm8 neurons could be produced during a longer temporal window than abundant neuronal types produced from larger domains20. However, more studies will be required to refine our understanding of how neuronal abundance is regulated.

Lastly, the intersection of a single temporal window, a single spatial origin, and a single Notch status has been presumed to encode a single neuronal identity (Fig. 1). However, as can be seen in Fig. 3c, in many cases such a combination encodes more than one neuronal type. Some cases are explained by the recent discovery of a fourth specification mechanism, in which neuroblasts from the same spatial, temporal and Notch origin but born at different developmental times produce different neuronal types18. Another possibility is that two neuronal types that appear to be produced from the same developmental origin could in fact be produced in different subdivisions of the same spatial domain, or from two different temporal windows that we were not able to distinguish. For instance, our results and others show that vOptix can in fact be subdivided into 4 subdomains: the anterior third producing pDm839, the middle third producing yDm839, and the posterior third split between an Optix+ Dpp− and an Optix+ Dpp+ domain (Fig. 3e). Moreover, although we had previously placed Tm1/2/4/6 in the Hth/Opa temporal window28, because marker genes did not allow us to distinguish these highly similar cell-types in the L3 main OPC, Zhang and colleagues41 have shown that Tm1 and Tm4 are generated in a first window of Opa expression while Tm2 and Tm6 are produced later, in a second window of Opa expression. Because we lack finer temporal window resolution for most main OPC clusters, we used the broad temporal windows we had previously characterized28 (Fig. 3c) for our analyses.

Regulation of terminal selector expression by specification factors

The expression of a given TS can theoretically be controlled by a single patterning mechanism, any combination of two patterning mechanisms, or all three patterning mechanisms (Fig. 4a). Moreover, each individual TS expression can either be regulated similarly in all neuronal types, or be controlled through different regulatory mechanisms across neuronal types, a process we previously described as “phenotypic convergence”12 (Fig. 4b). To characterize how TS expression could be regulated, we determined which of the TSCs we previously identified11,17 were expressed in all neuronal clusters that share one, two or three (Extended Data Fig. 7–9) “specification modules” (defined as “any temporal window, spatial origin, or Notch status that produces a neuronal type”). If the expression of a TSC is activated by a certain combination of specification modules, this TSC should be expressed in all neuronal types that share this combination (and potentially in other neuronal types, but through a different regulatory mechanism).

Fig. 4: Terminal selector candidate expression is correlated with developmental origin.

Fig. 4:

a) A given TS can be regulated by one, two, or three of the main OPC patterning mechanisms.

b) A given TS can be regulated similarly in all neuronal types (e.g. neurons 1 and 2), or through different mechanisms in different neurons (e.g., neurons 3 vs. 4).

a-b) N: Notch origin, S: spatial origin, T: temporal origin.

c) Only two TSCs have their expression perfectly correlated with a single specification module. hth (top) is continuously expressed in the main OPC clusters (x-axis) from the Hth temporal window, and ap (bottom) is continuously expressed in NotchON main OPC clusters. Clusters are grouped either according to their temporal window (top) or their Notch status (bottom). Gene expression is represented in light blue.

d) The expression of most TSCs is correlated with combinations of specification modules. For instance, ey is expressed in all Hth/Opa NotchOFF neurons, all Erm/Ey NotchOFF neurons, and all Ey/Hbn NotchOFF neurons that are not from the vDpp domain. A potential model of ey expression regulation is indicated (cluster 181 might also be the only Ey/Hbn cluster produced in both the vDpp and the dDpp domain, Fig. 2f). Clusters are grouped by combinations of one temporal and one Notch specification module, and gene expression is represented in light blue. Combinations of two specification modules containing a single cluster were not plotted.

These analyses identified two TSCs that could be under the control of a single patterning mechanism. First, apterous (ap, Fig. 4c), which was expected since neurons were identified to be NotchON based on ap expression26,28. Second, hth, which is continuously expressed exclusively in all neurons from the Hth temporal window (Fig. 4c, Extended Data Fig. 7). Although we did not identify any TSC continuously expressed exclusively in all neurons from any single spatial origin, TSCs specific to smaller spatial domains might still exist, but we could not identify them due to the incomplete resolution of our determination of spatial origins.

The expression of most TSCs was better correlated with combinations of all three patterning mechanisms, and the same TSC was often correlated with distinct combinations of specification modules in different neuronal types. For instance, ey is expressed continuously in NotchOFF neurons from both the Hth/Opa and Erm/Ey temporal windows (as a TSC in neurons, not as a tTF in neuroblasts) (Fig. 4d), regardless of their spatial origin, and from NotchOFF neurons from the Ey/Hbn temporal window except for those produced in the vDpp region (Fig. 3c). Therefore, ey expression seems to be controlled by at least three different combinations of specification modules (one potential scenario is shown in Fig. 4d). This is consistent with what was known for concentric genes, which are temporally expressed TSCs17,28,31 whose expression is further refined (excepted for hth, see above) by Notch signaling, such as bsh that is expressed only in the NotchON progeny of the Hth temporal window26, or spatial origin, as indicated by the regionalized expression pattern of some of them in the main OPC28.

Similar to what was presented above for ey, the heatmaps in Extended Data Fig. 7–9 can be used to form hypotheses on the specification modules controlling the expression of 66% of the TSCs (Extended Data Fig. 9b). We did not find correlations with developmental origin and the expression of 34% of TSCs, which could therefore be controlled by factors uncorrelated with the patterning mechanisms. However, these TSCs are very sparsely expressed (median of 3 neuronal clusters, Extended Data Fig. 9b): it is therefore very likely that improving spatial and temporal origin resolution would allow finding correlations between their expression and developmental origin. To illustrate the value of the correlations we identified, we present in Extended Data Fig. 10 already published26,28,29,40,41 experimental validation for 16 regulatory relationships that could have been hypothesized from our heatmaps (see also Methods on how to use these heatmaps).

Regulation of the presence of neuronal features by the terminal selectors

Next, we explored the regulation of gene expression by the TSCs, first by studying the regulation of the cholinergic marker VAChT, the glutamatergic marker VGlut, and the GABAergic marker VGAT in main OPC neurons. VAChT is expressed in 37 clusters, VGlut in 36 clusters, and VGAT in 24 clusters (Fig. 5a). If their expression were regulated by different TSC combinations in each of these neuronal types (model in Fig. 5b), this would require 37, 36 and 24 regulatory mechanisms, respectively. Strikingly, however, all neurons from the same temporal and Notch origin have the same neurotransmitter identity28 (Fig. 5a). VAChT is exclusively expressed in all NotchON neurons, except in the last temporal window where it is only expressed in NotchOFF neurons. VGAT is exclusively expressed in all NotchOFF neurons from the Hth, Hth/Opa, and Ey/Hbn temporal windows. VGlut is exclusively expressed in all NotchOFF neurons from the Opa/Erm, Erm/Ey, Hbn/Opa/Slp, and Slp/D temporal windows, and in all NotchON neurons from the D/Barh1 temporal window. This expression pattern is difficult to explain with a model in which neurotransmitter identity is regulated differently in each neuronal type (Fig. 5b).

Fig. 5: Correlation between developmental origin, terminal selector candidate expression, and effector gene expression.

Fig. 5:

a) Expression of the indicated genes (y-axis) in main OPC clusters (x-axis) grouped by combinations of one temporal and one Notch specification module. Light blue indicates either the clusters in which the TSCs are continuously expressed (top), or the clusters in which the markers of neurotransmitter identity are expressed at the adult stage (bottom).

b, c) Two models for the regulation of the expression of effector genes by TSs. Their expression could be controlled by different TSs in each neuronal type (b), or by the same TSs in all neuronal types sharing a given developmental origin (abbreviated dev. origin) (c).

d-f) Potential regulatory mechanisms for VAChT (d), VGAT (e), and VGlut (f) expression in the main OPC.

g) Expression of non-TF non-CSS genes (y-axis) in main OPC clusters (x-axis) grouped by combinations of one temporal and one Notch specification module. The names of the genes are indicated on the plot in Supplementary Data 2. Light blue: Expressed in this cluster and in at least 75% of clusters from the same combination of specification modules at the adult stage, Medium blue: Expressed in this cluster and in less than 75% of clusters from the same combination of specification modules at the adult stage, Dark blue: gene not expressed (Methods).

a, g) Combinations of two specification modules containing a single cluster were not plotted.

As described previously, it is also possible to find TSCs expressed in all neurons for each combination of temporal and Notch origins (Extended Data Fig. 8a). These TSCs could therefore be responsible for the regulation of VAChT, VGlut and VGAT (model in Fig. 5c), which would explain why their expression is correlated with the temporal and Notch origin of the neurons, and is more parsimonious than the previous model (Fig. 5b). This hypothesis also fits experimental evidence. We have shown that the TSC ap is an activator of cholinergic identity in the optic lobe12, and it is expressed in all NotchON neurons26. Therefore, VAChT expression could be activated by ap in all NotchON neurons, except in the D/BarH1 temporal window. In this window, some TSCs common to all D/BarH1 NotchON neurons could prevent the activation of VAChT expression despite the presence of ap, while some TSCs common to all D/BarH1 NotchOFF neurons could activate VAChT expression despite the absence of ap (Fig. 4d, Extended Data Fig. 8A). Moreover, we have shown that Lim3 is an activator of GABAergic identity12. It is expressed in GABAergic neurons, which come from only 3 developmental origins (Fig. 5a). However, Lim3 is also expressed in non-GABAergic neurons from two other developmental origins (Opa/Erm NotchOFF and Erm/Ey NotchOFF neurons, Fig. 5a), where the activation of VGAT expression by Lim3 could be prevented by the TSCs specific to these two origins (Fig. 5e, Extended Data Fig. 8a). Lastly, we have shown that tj is an activator of glutamatergic identity12. It is exclusively expressed in Hbn/Opa/Slp NotchOFF neurons, which are glutamatergic (Fig. 5a), making it is likely that VGlut expression is activated by tj in neurons from this developmental origin. VGlut is also expressed in neurons that do not express tj and that come from 4 other developmental origins. VGlut expression cannot be activated by tj in these 4 groups of neurons, but could be regulated in each of these 4 groups by the TSCs that are specific to these groups (Fig. 5f, Extended Data Fig. 8a).

We then hypothesized that other genes or even morphological features could be regulated in a similar manner, i.e. by the same TSCs in all neurons from the same developmental origin (Fig. 5c). To test this hypothesis, we surveyed the Virtual Fly Brain42 website as well as more than 40 publications to build Supplementary Table 2 that summarizes about 600 morphological features for the more than 70 neuronal types for which we have identified the corresponding clusters in our scRNA-seq atlas. This table can be amended with newly identified clusters, or new characterizations of neuronal morphologies9,10, and used as the basis for further high-throughput studies of the regulation of neuronal morphology. This table, as well as binarized gene expression data, allowed us to identify genes and morphological features present in all neuronal types sharing any combination of one to three specification modules at each developmental stage, consistent with their regulation by TSCs associated with the corresponding developmental origin (Fig. 5c). Fig. 5g presents examples of genes expressed at the adult stage in neurons sharing the same combination of temporal and Notch origin, to illustrate that many genes follow the same expression pattern as the markers of neurotransmitter identity and therefore may be under the control of the same TSCs. All other plots are provided in Supplementary Data 2 for gene expression (split into TFs, Cell Surface and Secreted proteins – CSSs, and non-TFs/non-CSSs for readability) and in Supplementary Fig. 1, 2 for morphological features. Many correlations between developmental origin and morphological features can be observed and are consistent with our work that showed in an independent manner that developmental origin is correlated with neuropil targeting depth, synapse location, and even visual functions of main OPC neurons43.

Our results show that neurons sharing combinations of specification modules share the expression of some TSCs and some modules of neuronal features. These TSCs are therefore good candidates to regulate their associated module of neuronal features, which was validated12 for ap, tj and Lim3 in the case of neurotransmitter identity. In such cases, TSCs would regulate the same targets in all neurons. This and other results17 are consistent with a model whereby it is not required for all TSs to be involved in the regulation of all terminal effectors in a given neuron, but rather, any effector gene can evolve to be regulated by any combination of available TSs in that neuron. However, for most differentially expressed genes, their expression was not associated with any combination of specification modules. These genes could therefore be regulated by any of the TSs expressed in a given cell type. Further studies will be required to understand how the expression of these genes is regulated.

Specification modules and terminal selectors encode similar information

According to the TS model, in which a TS code maintains the neuronal identity established in progenitors, TS expression in neurons must be sufficient to fully encode the information inherited from the specification modules in the progenitors, which we tested in a high-throughput manner using machine learning. Intuitively, if several neuronal types share a given morphological feature or express a given gene, and all specifically express a TF, this TF is a good candidate regulator of this feature. Its expression could therefore be used to predict the presence of the feature in other neuronal types. To identify such correlations for all neuronal features we used the random forest machine learning algorithm44 (Fig. 6a). For each neuronal feature at any given developmental stage from P15 to adult, we trained a random forest model on a randomly sampled subset of clusters (“training set”). This established a model that correlates the expression of potential regulators with the presence or absence of a given neuronal feature. These potential regulators, called predictors, were either our previously published TSCs11,17, TFs (including TSCs), CSSs, or the specification modules, and the neuronal features predicted were either the expression of a gene or the presence of a morphological feature (Supplementary Table 2). To assess the quality of each model, we used them to predict the expression of the feature they were trained to predict in a “test set”, i.e., neuronal clusters not used to train the model and for which we already knew the real status of the feature. The similarity between predicted and actual status of the feature was then evaluated by using the Matthews Correlation Coefficient (MCC, a value of 1 denotes perfect correlation and -1 perfect anti-correlation).

Fig. 6: Specification modules and terminal selector candidates encode similar information.

Fig. 6:

a) Machine learning approach used to find candidate regulators (TFs, TSCs, CSSs, or specification modules) for neuronal features (gene expression or morphologies). Clus. = clusters, Pred. = predictor, RFM = Random Forest Model, MCC = Matthews Correlation Coefficient.

b-d) Matthews Correlation Coefficient between the observed expression of TSCs (b) or marker genes (c, d), and their expression predicted using either specific TFs (i.e., all TFs differentially expressed between neuronal types, b-d), specification modules (b, c), or TSCs (d). Each gray dot represents the MCC value for a given gene, and the values are summarized as boxplots that display the first, second and third quartiles. Whiskers extend from the box to the highest or lowest values in the 1.5 interquartile range, and outlying data points are represented by large black dots.

We first built models using specification modules to predict TSC expression at all stages, and obtained high predictivity (median MCC around 0.6, Fig. 6b). Considering the imperfect resolution of our determination of both temporal and spatial origins, this suggests that TSC expression in neurons is highly correlated with specification modules, consistent with the previous sections. However, developmental origin performed worse when predicting the expression of TFs other than TSCs (median MCC around 0.3, Supplementary Fig. 3a), suggesting that the expression of these TFs is not directly or not only controlled by temporal, spatial and Notch origins.

Second, we used either specification modules or TSC expression to predict the expression of marker genes (i.e., differentially expressed genes that are therefore the most representative of each neuronal identity). We compared the results to those obtained using all differentially expressed TFs as predictors, which include TSCs, and which we call “specific TFs”, because they should produce the best possible models of marker gene expression. Both TSC expression and developmental origin performed similarly to using specific TFs to predict marker gene expression, and sometimes performed even better (Fig. 6c, d). As expected, when using spatial, temporal, or Notch origin independently, the predictivity was lower than when using all three combined (Supplementary Fig. 3b). Similar results were obtained when predicting morphological features (Supplementary Fig. 3c, d), although only 12–60 features could be tested (see Methods), compared to hundreds of marker genes. These results show that specification modules (n=27) and TSCs (71<n<74 depending on developmental stage) can predict neuronal identity with a similar power as all differentially expressed TFs (222<n<410) despite their much lower number. When predicting the expression of genes that are neither marker genes nor pan-neuronal (i.e., expressed in many but not all neurons), however, the performance was slightly lower for TSCs, and much lower for developmental origin (Supplementary Fig. 3e, f). This could be because the regulation of their expression is more complex than the one of marker genes: for instance, more broadly expressed genes may be more likely to undergo phenotypic convergence12 (i.e., to be regulated differently in different neuronal types) because they are expressed in neurons from a wider variety of developmental origins.

Lastly, it has been suggested that the information encoded by the spatial factors could be maintained epigenetically in neurons, likely by differential chromatin accessibility 22,45. This entails that TFs would have different targets based on the spatial origin of the neurons. In this case, our random forest models based on TSC or TF expression would not work due to lack of information. However, using spatial origin (in fact, any specification module) as predictors in addition to TF or TSC expression did not lead to increased predictivity of our random forest models for marker gene expression (Supplementary Fig. 4a, b). This suggests that developmental origin does not contain information that is not already encoded into the neuronal TSC combinations, although these results could also be explained by an insufficient precision in our characterization of the developmental origin of each neuronal type.

Each trained random forest model provides a ranking of the predictors according to how important they are for the performance of the model. If a model performs well, it has identified correlations present both in the training and the test set, and therefore its best ranked predictors are candidate regulators of the feature of interest. For instance, the random forest model established for Gad1 using all TFs in the main OPC identified Lim3, an experimentally validated activator of Gad1 expression12, as second-best candidate regulator. This model obtained an MCC of 0.745, suggesting that models above this MCC (432 genes at the adult stage) can be considered as yielding high quality candidates. In addition to using TF expression to predict gene expression or the presence of morphological features (Supplementary Fig. 5a, b), we also produced models predicting the presence of morphological features from CSS expression (Supplementary Fig. 5c). Although it is out of the scope of this study, these models can be used to test experimentally the regulation of biologically relevant features, and we provide in Supplementary Data 3 tables listing the candidate regulators identified for each neuronal feature. However, because random forest can only identify correlations, using additional approaches to reduce the number of candidate regulators for a given feature is advised (see Methods).

Discussion

In this work, we identified the spatial origins of the main OPC neuronal types and explored how patterning information establishes TS expression and ultimately neuronal features. Our results suggest that the expression of a given TS can be controlled through multiple regulatory mechanisms, each activated by a different combination of one to three specification modules. This entails that the TS combination specific to a given neuronal type is established by integrating different subsets of the patterning information for different TSs (Fig. 7). Although the correlations we identified between specification modules and TSC expression could be used as a basis to identify motifs bound by specification factors in the enhancers of TSCs, further work is required to validate these correlations and to characterize the corresponding molecular mechanisms. Notably, in some cases, the expression of a TS could be regulated indirectly by a specification factor, and some TSs have also been shown to cross-regulate each other17,46. In addition, because spatial factors are expressed in the neuroepithelium but not in the neuroblasts, temporal TFs are expressed in neuroblasts but often not maintained in neurons, and Notch effectors are only transiently expressed in newborn neurons, it will be of interest to understand when exactly the TS combinations are established. This could be an iterative process, in which for instance epigenetic marks (e.g., chromatin accessibility) are established in the neuroepithelium by the spatial factors, these epigenetic marks interact with tTFs and produce preliminary TS combinations in neuroblasts and/or ganglion mother cells, which are then finalized by Notch signaling in newborn neurons. It could also happen in a single step in newborn neurons, provided that spatial and temporal origins are remembered up to this point, for instance by the expression of intermediate factors or by epigenetic mechanisms. Moreover, our work provides testable hypotheses on how TS expression controls neuronal type production. For instance, Mi1 is produced in both Optix and Vsx, while Mi9 is produced in Optix but not Vsx (Fig. 3c). This could be explained if the expression of all Mi1-specific TSs was activated (or not repressed) by both Optix and Vsx, while the expression of at least one of the Mi9-specific TS was either activated by Optix or repressed by Vsx.

Fig. 7: Establishment and maintenance of neuronal features in the Drosophila optic lobe.

Fig. 7:

Each neuronal type (grey box) expresses a unique combination of TSs (TSa-n, top part of the box). The expression of each of these TSs is established either by one, two, or three specification modules (spatial origin, temporal window, Notch signaling). Moreover, some effector genes (in colors, middle part of the box) are regulated by the TSs corresponding to a given combination of specification modules (ex: a given Notch status activates the expression of TSd, which activates the expression of two effector genes in all neurons sharing this Notch status). The expression of each effector gene represented in grey (bottom part of the box) is not correlated with any combination of specification modules, and can be regulated by any combination of the TSs expressed in this neuronal type.

Our results also have implications for neuronal type evolution. Indeed, new neurons could appear following the emergence of a new tTF, or of a new spatial specification factor. However, these large-scale modifications of neuronal specification could have deleterious effects, either by leading to the appearance of many new neuronal types at once, or by affecting existing specification mechanisms. For instance, the expression of a new tTF would likely overlap with the expression of existing tTFs, therefore affecting existing temporal windows in addition to creating a new one. However, the discovery of additional spatial subdivisions in this work and others20,39, the recent discovery of additional broad temporal windows that could be further split into narrower ones 28,29, and the discovery of a fourth patterning mechanism in the optic lobe18, suggest that the specification factors have the potential to encode many more neuronal types than the ones produced in the optic lobe. New neuronal types could thus emerge either by rescuing a type previously eliminated by cell death, or for instance by splitting a spatial domain within a single temporal window (e.g., if a single TS expressed in a neuronal type produced in vOptix/vVsx became repressed by Optix, this would lead to the production of two neuronal types: one from vOptix and one from vVsx). Because any effector gene can evolve to be regulated by any combination of available TS, most of the genes expressed in a newly evolved neuronal type cannot be predicted. Nevertheless, according to our model (Fig. 7), a newly produced neuronal type from a given temporal window and of a given Notch status should present a pre-determined module of molecular and morphological features associated with its temporal+Notch origin, such as VAChT expression in all NotchON neurons except in the last temporal window (Fig. 5a). In addition, such a neuronal type would also present modules of features associated with any combination of its specific temporal, spatial and Notch origins. These pre-determined modules of features could bias newly evolved neurons to integrate into given circuits or to assume given functions: similar environments could lead to the selection of new neuronal types with similar functions, and therefore produced according to similar specification mechanisms. Such reasoning could be tested experimentally.

Methods

Immunohistochemistry

Whole brains were dissected in Schneider’s media (Sigma S0146), transferred to Schneider’s media on ice for no more than 30 min, and fixed using 4% paraformaldehyde (Electron Microscopy Sciences 15710) diluted in PBS 1X, for 20 min at room temperature. They were then rinsed 3 times in 0.5% Triton™ X-100 (Sigma) diluted in PBS 1X (PBTx), washed for 30 min in PBTx, and incubated at least 30 min in PBTx with 5% horse (ThermoFisher 26050070) or goat (ThermoFisher 16210064) serum (PBTx-block) at room temperature. They were then incubated at 4°C for 1-2 overnights in primary antibodies diluted in PBTx-block. They were then rinsed 3 times in PBTx, incubated two times for 30 min in PBTx, and incubated at 4°C overnight in secondary antibodies diluted in PBTx-block. Lastly, they were rinsed 3 times in PBTx, incubated two times for 30 min in PBTx, and mounted in SlowFade™ (ThermoFisher S36936) before imaging with a Leica SP8 confocal microscope using a 63x glycerol objective. Images were processed in Fiji and Adobe Illustrator.

The following primary antibodies were used: mouse anti-Svp (1:20) (DSHB), mouse anti-tup (1:100) (DSHB), mouse anti-Dac 2-3 (1:40) (DSHB), mouse anti-Cut (1:10) (DSHB), rabbit anti-Dichaete (1:500) (modENCODE), rabbit anti-toy (1:500)11, rabbit anti-Bsh (1:1800) 11, rat anti-vvl (1:50)28, guinea pig anti-Vsx1 (1:100)17, guinea pig anti-Hth (1:500) (gift from Richard Mann), guinea pig anti-TfAP-2 (1:100)28, rat anti-drgx (1:250)17, rat anti-erm (1:100) 28, guinea pig anti-run (1:500)28, rat anti-Ey (1:100) (this study), guinea pig anti-Kn (1:100)28, rabbit anti-SoxN (1:250)17, guinea pig anti-Tj (1:250) (gift from Dorothea Godt47), rabbit anti-Dll (1:100)28, chicken anti-GFP (1:200) (Millipore Sigma), rat anti-Ecad (1:200) (DSHB).

The following secondary antibodies, all from donkey, were used: anti-rat DyLight 405 (1:100), anti-rabbit DyLight 405 (1:100), anti-guinea pig DyLight 405 (1:100), anti-mouse Alexa Fluor 488 (1:400), anti-mouse Alexa Fluor 555 (1:400), anti-guinea pig Cy3 (1:400), anti-rabbit Alexa Fluor 555 (1:400), anti-rat Alexa Fluor 647 (1:200), anti-rabbit Alexa Fluor 647 (1:200), and anti-guinea pig Alexa Fluor 647 (1:200).

Antibody generation

A rat polyclonal antibody against Ey was generated by Genscript using the following epitope (https://www.genscript.com/):

MFTLQPTPTAIGTVVPPWSAGTLIERLPSLEDMAHKDNVIAMRNLPCLGTAGGSGLGGIAGKPSPTMEAVEASTASHPHSTSSYFATTYYHLTDDECHSGVNQLGGVFVGGRPLPDSTRQKIVELAHSGARPCDISRILQVSNGCVSKILGRYYETGSIRPRAIGGSKPRVATAEVVSKISQYKRECPSIFAWEIRDRLLQENVCTNDNIPSVSSINRVLRNLAAQKEQQSTGSGSSSTSAGNSISAKVSVSIGGNVSNVASGSRGTLSSSTDLMQTATPLNSSESGGASNSGEGSEQEAIYEKLRLLNTQHAAGPGPLEPARAAPLVGQSPNHLGTRSSHPQLVHGNHQALQQHQQQSWPPRHYSGSWYPTSLSEIPISSAPNIASVTAYASGPSLAHSLSPPNDIESLASIGHQRNCPVATEDIHLKKELDGHQSDETGSGEGENSNGGASNIGNTEDDQARLILKRKLQRNRTSFTNDQIDSLEKEFERTHYPDVFA

Determination of the spatial origin of main OPC clusters

Obtaining the animals for scRNA-seq and testing GFP expression

The genotypes of the Drosophila lines used were (MC stands for memory cassette):

  • Optix-ts-MC line: 20xUAS-FlpG5::PEST; Optix-Gal4,tub-Gal80ts/CyO; ubi>STOP>Stinger/TM6

  • vOptix-MC line: w*,disco-T2A-VP16AD; optix-T2A-Gal4DBD/+; UAS-Flp.D/UAS-Flp.Exel,UAS-RedStinger,Ubi>Stop>Stinger(G-Trace)

  • dOptix>Gal4: w1118; salr-T2A-VP16AD,Optix-T2A-Gal4DBD/CyO

  • hh-ts-MC line: 20xUAS-FlpG5::PEST; ubi>STOP>Stinger/CyO; hh-Gal4,tub-Gal80ts/TM6B,Tb

  • pxb-MC line: 20xUAS-FlpG5::PEST; ubi>STOP>Stinger/CyO-nlsGFP; pxb-T2A-Gal4/TM6B,Tb

  • dpp>GFP line: w; (UAS-Stinger; dpp-Gal4)/[SM6a:TM6B, Tb, actGFP]

For the Optix dataset, Canton S virgins crossed with Optix-ts-MC males were allowed to lay eggs for 2 hours at 18°C. The eggs were then placed at 29°C for 2 days and 9 hours (late L2 stage), then back at 18°C for the remainder of their development. GFP expression was tested in L3 wandering larvae and sequencing was performed in adults (Fig. 2b). We produced 3 libraries from the same cell suspension of optic lobes from 5 female flies that were a few hours old.

For the vOptix dataset, vOptix-MC flies were kept at 25°C. GFP expression was tested in L3 wandering larvae and sequencing performed in adults (Fig. 2b). We produced 2 libraries from the same cell suspension of the optic lobes of 15 female flies that were more than two-week-old. Some ectopic expression in the main OPC was observed in this line (Extended Data Fig. 1c, d), which we accounted for in later analyses.

For the dOptix dataset, w; ; UAS-Stinger virgins were crossed with dOptix>Gal4 males and kept at 25°C. GFP expression was tested, and sequencing was performed, in P0-P12.5 pupae (Fig. 2b). We produced 2 libraries from the same cell suspension of 24 optic lobes.

For the hh dataset, Canton S virgins crossed with hh-ts-MC males were allowed to lay eggs for 2 hours at 18°C. The eggs were then placed at 29°C for 2 days and 3 to 9 hours (mid to late L2 stage), then back at 18°C for the reminder of their development. GFP expression was tested in L3 wandering larvae and sequencing was performed in adults (Fig. 2b). We produced simultaneously 2 libraries from independently produced cell suspensions, each from the optic lobes of 7 two-week-old female flies.

For the pxb dataset, Canton S virgins were crossed with pxb-MC males and kept at 25°C.

GFP expression was tested, and sequencing was performed, in P0-P12.5 pupae (Fig. 2b). We produced 1 library from 8 optic lobes. Some ectopic expression in the main OPC was observed (Extended Data Fig. 1e). We therefore checked the pattern of GFP expression of each optic lobe by epifluorescence microscopy before sorting and sequencing. However, due to the low resolution of this technique, it is conceivable that small amounts of false positive cells were not detected.

For the dpp dataset, Canton S virgins were crossed with dpp>GFP males and kept at 25°C. GFP expression was tested, and sequencing was performed, in P0-P12.5 pupae (Fig. 2b). We produced simultaneously 2 libraries from independently produced cell suspensions, each from 12 optic lobes of young pupae.

All animals used for scRNA-seq were selected for the lack of balancer chromosomes. P0-P12.5 animals were obtained by selecting pupae in which the head had not yet detached from the pupal case47.

For the wg dataset, the sequencing of the wg-Gal4 line is described in El-Danaf et al. 202421.

Sample preparation and single-cell mRNA-sequencing

We followed a similar cell dissociation protocol to generate all spatial domain datasets. Whole brains were extracted in ice-cold Schneider’s media. The brains were then transferred into DPBS 1X without calcium and magnesium (Corning/Fisher 21-031-CV). We used an epifluorescence microscope to remove brains with ectopic main OPC GFP expression (see main text), and placed the brains back in Schneider’s media. The brains were then incubated at 25°C in Schneider’s media with 2 mg/mL Collagenase (Sigma C0130) and 2 mg/mL Dispase (Sigma D4693-1G), for 1 hr 30 min (adult brains) or for 30 min (pupal brains), then washed 3 times with ice-cold Schneider’s media to stop the dissociation. Optic lobes were then separated from the central brain, transferred to ice-cold DPBS + 0.04% BSA (ThermoFisher AM2618), and washed twice with ice-cold DPBS + 0.04% BSA. For pupal brains, the separation between central brain and optic lobes can be seen as a difference in texture: the optic lobe forms a smoother and cohesive structure that can be removed using a fine tungsten needle. Optic lobes and 150 μL of DPBS + 0.04% BSA were then transferred to a Eppendorf™ DNA LoBind™ Tubes (Fisher Scientific 13-698-791), and dissociated by pipetting vigorously 50 times, resting the cells for 1 minute on ice, and pipetting 50 times more. The cell suspension was then checked under a dissecting microscope, and we pipetted up to 50 more times if the size of remaining clumps was too big. We did not try to dissociate small clumps since they are most likely neuropils, not somas, and indeed we empirically noticed a better quality of the data obtained after lowering pipetting repetitions. The suspension was then passed through a 20 μm strainer (pluriSelect 43-10020-40), and the filter was washed with 400 μL of DPBS + 0.04% BSA which we also collected (pipetting from the opposite side may be necessary to collect the entire cell suspension).

Stinger(nuclear GFP)-positive cells were sorted with a FACSAria II using FACSDiva 8.0.2 to set the gates shown in Supplementary Data 1, and we loaded the single-cell sequencing chip with 43.5 μL of the cell suspension provided by the FACS, undiluted, mixed with the master mix. Moreover, we used a tip coated with 70 μL of DPBS + 0.04% BSA and 0.02% Tween-20 (Bio-Rad 1662404) to sample the cell suspension, and kept the same tip to load the chip, which we empirically noticed drastically reduces the loss of cells. All steps of single-cell sequencing were performed following the Chromium Next GEM Single Cell 3′ Reagent Kits v3.1 (Dual Index) instructions.

Sequencing was performed using Illumina NextSeq 500 or Illumina NovaSeq 6000 at the Genomics Core at the Center for Genomics and Systems Biology at New York University. The data was mapped to the D. melanogaster genome assembly BDGP6.88 using cell ranger 6.0.1.

Supervised annotation of the scRNA-seq datasets:

The libraries were analyzed using Seurat 4.3.0. For each library, a Seurat Object was created with all genes expressed at least in 3 cells, and all cells expressing at least 200 genes. The objects were then filtered by keeping all cells below a certain percentage of mitochondrial genes, below a certain number of UMIs, and above a certain number of genes, chosen according to the distribution of these parameters (Supplementary Data 1). The thresholds chosen were identical for all libraries acquired at a given day with flies of a given genotype, but were different otherwise. For the Optix libraries the thresholds were 7/17000/800 (percent of mitochondrial genes, number of UMIs, number of genes), for the vOptix libraries 10/10000/500, for the dOptix libraries 5/20000/1000, for the hh libraries 10/20000/700, for the dpp libraries 5/20000/900, for the pxb library 5/30000/1300, for the wg libraries 10/25000/1200. After filtering, we obtained an Optix dataset of 19,036 cells (3 libraries), a hh dataset of 9,182 cells (2 libraries), a vOptix dataset of 10,596 cells (2 libraries), a dOptix dataset of 12,302 cells (2 libraries), a pxb dataset of 4,988 cells (1 library), a dpp dataset of 12,862 cells (2 libraries), and a wg dataset of 12,974 cells (2 libraries). For each library we then ran NormalizeData, FindVariableFeatures, and ScaleData with default parameters, as well as RunPCA, RunTSNE, and RunUMAP with default parameters and a dimensionality of 150. Dimensionality reductions were run purely for visualization purposes; they were not used to perform any clustering. Instead, we used the normalized expression of marker genes and our neural network classifier11 to assign each cell of each library from this study to its corresponding cluster in our previously published scRNA-seq atlas11 (metadata fields “NN_Cluster_Number”), and to give this assignment a confidence score between 0 and 1 (metadata field “Confidence_NN_Cluster_Number”). Notably, since we know which clusters are produced in the main OPC28, this allowed us to identify non-main OPC cells present in our datasets, including the ones due to expression of our reporters outside of the main OPC. For all figures, we replaced cluster numbers by cluster annotations (metadata fields “Annotation”) using Supplementary Table 3. Lastly, if at least 80% of the cells of a given Class were annotated with a confidence score strictly below 0.5, the Class was flagged as “low confidence” by assigning it the value “1” in the metadata field “cells_flagged”, and its annotation was changed from “X” to ““LowC_X”.

Binarization of the spatial origin of each main OPC cluster

We then normalized the abundance of each neuronal Class in each dataset. Indeed, a neuron produced exclusively in the Vsx domain is expected to be relatively more abundant in a library produced from sorted pxb cells compared to our published atlas, which contains neurons produced in all domains. However, a neuron produced in each domain equally should have a similar abundance in all datasets (Extended Data Fig. 2b). To allow comparisons of Class abundances between libraries, we used the abundances of neurons produced in the whole main OPC, since they represent a constant between main OPC domains. In each library, we therefore divided the abundance of each Class by the abundance of all the Classes corresponding to neuronal clusters produced in all the main OPC. We also normalized the abundance of each Class in our single cell atlas but averaged the normalized abundance across all libraries to simplify the comparisons. We performed a first round of normalization using Mi1, Tm1 and T1. We then used the results of this round of normalization to determine that T1, Mi1, Tm1, Tm2, Tm4, and cluster 138 were produced in all main OPC domains and therefore used them for a second round of normalization. Because this second round of normalization did not show additional neurons produced in the whole main OPC, we did not perform further rounds of normalization. However, these clusters are not produced in the wg domain, so they could not be used to normalize wg libraries. We therefore only normalized this dataset by dividing the abundance of each Class by the abundance of its most abundant cluster.

However, normalized abundances are not enough to decide whether a neuronal type is produced in a given domain, because rare cell types would still represent only a small part of any dataset (for instance cluster 38 is produced in the vVsx domain, which represents only half of the Vsx domain and thus the pxb dataset). We therefore computed an enrichment score for each Class in each spatial domain (Extended Data Fig. 2b). To do so, we divided the average normalized abundance of a Class in all libraries of a spatial domain dataset by its average normalized abundance in our single cell atlas. Theoretically, any score above 1 should indicate an enrichment of the Class in a main OPC domain compared to the whole main OPC, and therefore indicate that this cluster is produced in this domain. However, for a given Class, the minimal and maximal normalized abundances across libraries of a same main OPC domain are usually variable (see Supplementary Data 1), probably due to technical reasons (dissociation, FACS…), and therefore enrichment scores could not be analyzed in such a straightforward manner.

Therefore, we then binarized these “enrichment scores” by choosing thresholds based on our experimental validations of spatial origin (Extended Data Fig. 4). For the Optix dataset, we chose Dm4 as a threshold because this is the least abundant neuronal type that we ascertained to come from this domain. For the dOptix dataset, we chose Tm1 as a threshold instead of TmY5a or Tm20 (that we know are produced in the dOptix domain), because some of these might have been annotated as immature neurons (with which they form a trajectory on UMAP visualization, Extended Data Fig. 6). For the vOptix dataset, we chose Mi9 as a threshold instead of cluster 161, because cluster 161 is a very small cluster that is unlikely to be produced in all of the vVsx, vOptix and vDpp domains (which is where we found the expression of the markers for cluster 161, 167, 168, and Dm10, Supplementary Table 1). For the pxb and dpp datasets, we chose Tm1 as a threshold because out of all the neuronal types we experimentally validated as coming from these domains, Tm1 was the least abundant in both the pxb and dpp datasets. For the hh dataset, we chose cluster 38 as a threshold because we know it is produced exclusively in the ventral OPC. We did not choose Tm5ab because 1) this cluster contains two cell types, and 2) their markers (shared with other clusters) are found in the d/vVsx, dOptix and dDpp: the ventral main OPC represent only a small part of their potential domain of production.

We performed all these analyses using both our scRNA-seq atlas at the P15 and adult stages as a reference, because some clusters (for instance the TE cells) are only found at one of the two stages. However, we favored results obtained with the adult dataset, because it contains more cells, and at this stage their identities are more defined (there are no more immature neurons). Moreover, we performed these analyses for both main OPC and non-main OPC clusters but displayed these clusters separately in our plots. The presence of non-main OPC neurons in our datasets can mean that 1) the neuron is produced outside the main OPC but still expressed our reporters, 2) the neuron is in fact produced in the main OPC and was mistakenly assigned as non-main OPC neurons28.

Final assignment of a spatial origin to each main OPC cluster

We primarily used binarized enrichment to assign an origin to each Class. However, we deemed the binarized enrichment of a Class in a given dataset as less trustworthy if a class was flagged as being annotated with low confidence, and/or there were 3 cells or fewer of this class found in the dataset, and/or its enrichment value was +/− 0.05 compared to the binarization threshold (Extended Data Fig. 5). In this case, we also considered additional parameters (Supplementary Table 1), such as its localization on the UMAP (Extended Data Fig. 6). For instance, if “ClassA” contains only a few cells annotated with low confidence, and its cells are mixed on the UMAP with cells from “ClassB” that are annotated with high confidence, it is likely that ClassA cells are in fact ClassB cells misannotated by the neural network. We also accounted for the expression of vOptix-MC in the vVsx stripe and in the tips of the OPC: if a neuronal type is produced only in the Dpp domain, for instance, it should be absent in the Optix dataset but present in the vOptix dataset. Similar reasoning applies to a neuron produced in vVsx. Our choices are explained in Supplementary Table 1. Lastly, it should be noted that in the absence of dVsx and dDpp datasets, assignment to these domains could only be extrapolated. For instance, if a Class was found in the dpp and the hh datasets, it is clearly produced in vDpp, and maybe in dDpp. However, if the Class was also found in the vOptix and dOptix datasets, then the corresponding neuron is very likely to be also produced in the dDpp domain. Such determinations are also explained in Supplementary Table 1.

Identifying clusters corresponding to synperiodic neurons

Synperiodic neurons are the most abundant neurons in the optic lobe: they are present in ~800 copies, i.e., one per column. The clusters corresponding to synperiodic neurons should therefore be the largest ones in our scRNA-seq atlas11. Based on electron microscopy data10, among the clusters containing more than 1,000 cells in our adult scRNA-seq dataset, most of them clearly correspond to synperiodic neurons: Tm3 (1037 cells. Tm3 was previously described as ultra-periodic, i.e., being present in more than one copy per column7), T1 (892 cells), Tm1 (890 cells), Mi4 (889 cells), Mi9 (889 cells), Mi1 (887 cells), Tm9 (887 cells), Tm2 (883 cells), Tm20 (876 cells), Tm4 (833 cells), Tm6 (770 cells), Dm2 (726 cells). A few correspond to neurons of intermediate abundance: TmY5a (678 cells), Mi15 (582 cells), Dm8 (550 cells), TmY3 (409 cells), Dm10 (311 cells). The status of Dm3 is unclear, as our dataset contains only two corresponding clusters, Dm3a and Dm3b, although three abundant Dm3 subtypes have now been described (Dm3a, 601 cells, Dm3b, 571 cells, Dm3c, 401 cells). The list of neurons with quantified abundance did not include TmY8 and Tm2510, and two of our clusters containing more than 1,000 cells are unannotated (clusters 82 and 36).

Producing the table of morphological features

To produce the table, we surveyed the Virtual Fly Brain website (https://v2.virtualflybrain.org/) as well as 40 publications (Supplementary Table 2). The definition of each feature is indicated in the column “Comments” and the line corresponding to the feature in Supplementary Table 2. The confidence in the values are color coded as explained in the “legend” tab. One important characteristic of the table is that, for each feature, we distinguished between known absence in a cluster, or whether its presence was simply not evaluated, which we denoted by “NA” (in which case the cluster was removed when building machine learning models for this feature). Moreover, the control of neuronal morphology by TF or CSS expression is likely to be context specific. For instance, a CSS might be important to target lobula layer 4 in some neurons but have a completely different role in a neuron that does not enter the lobula. Therefore, some of the molecular features in Supplementary Table 2 are duplicated: in one case they are evaluated in all neuronal types, and in a second case (identified by adding “filtered” to the name of the feature) neurons that cannot express the feature (ex: targeting of lobula layer 4 for non-lobula neurons) are flagged with “NA”, for “non-applicable”, which also allowed us to remove them when building random forest models for these features.

Among the morphological features, layer targeting was one of the least straightforward to assess because the available information was often conflicted between different sources. This can be due to differences in layer definitions between different authors (for instance differences in the definition of the 4th medulla layer1,49), or different imaging techniques (for instance, Takemura and colleagues7 note that “In some cases […], the cells reconstructed from EM appear to have fewer processes than the light microscopy images. This relative sparsening results from difficulties in connecting the fine processes of the arborizations to the main body of the neuron during EM reconstruction.”). Moreover, the binarization of layer targeting data was difficult due to lack of defined threshold and inter-individual variation. For instance, Takemura and colleagues7 present two reconstructed Mi4 cells: one terminates in M9 while contacting the M10 layer border, and the other terminates in M10. One shows clear branching in M9, the other does not. And both show one arborization in M7, but these are much smaller than arborizations in other layers, and Mi4 is usually not reported to arborize in M7. We did our best to indicate low confidence in our assessment of targeting determinations by a color code (Supplementary Table 2), and generally chose a non-conservative approach in which for any each neuronal type we indicated all layers in which it was shown to produce significant arborizations, even if only one report of such instance was known. In the future, using synapse coordinates instead of layer targeting would be a more rigorous approach.

For connectivity in the lamina, we used Supplementary Table 2 of the study by Rivera-Alba and colleagues5, in which we removed all synapses that were considered as uncertain and grouped the cell types L4, L4+x and L4+y. We then converted the number of presynapses from cell type A to cell type B into a proportion of all presynapses from cell type A (number of presynapses made by A to B / total number of presynapses made by A). We did the same for the postsynapses (number of postsynapses made by A from B / total number of postsynapses made by A). For connectivity in the medulla, we utilized synaptic connections between neurons from the Fib25 dataset from neuPrint (GitHub Repository: connectome-neuprint/neuPrint; commit hash: 69fbda4). Briefly, we retrieved all annotated neurons that have any synapses, and we matched the annotated pre- and post-synapses with a confidence score to produce a table called “raw_connection_by_cell.csv” that contains all occurrences of a synapse between a given couple of pre and postsynaptic neurons. We kept only synapses with a perfect confidence score, and duplicated the results obtained for Tm9 into Tm9v and Tm9d (which assumes they connect to the same cell types). We then produced “Table_presynapses_to.csv” that contains the proportion of presynapses that a given cell type (ex: L2, column) makes with any other cell type (ex: Mi1, row), and “Table_postsynapses_by.csv” that contains the proportion of postsynapses that a given cell type (ex: L2, column) makes with any other cell type of the study (ex: Mi1, row). The code used for this re-analysis is provided in Supplementary Code 1. Both for connectivity in the lamina and the medulla, we also produced “Synapsed_to_A” and “Synapsed_by_A” features by binarizing the proportions of synapses by neuronal type A, or to neuronal type A, with a threshold of 5%.

In addition to features related to the shape and the connectivity of individual neuronal types, we also added a few “identity” features. For instance, the feature “Dm_identity_filtered” is true for all clusters corresponding to a Dm neuronal type (Dm1, Dm2, etc.), false for all clusters corresponding to non-Dm neuronal types, and “NA” (i.e., not used in machine learning) if the cluster is not annotated. Therefore, this feature was used to find good predictors of “being a Dm”. Similar identity features were done for other groups of neuronal types (Tms, Mis, etc.).

Production and evaluation of the random forest models

Selection of the clusters used

For each stage, we performed the analyses using either “All clusters” or “main OPC clusters”. By “All clusters”, we mean all neuronal clusters (as indicated in Supplementary Table 3), except a) clusters containing less than 5 cells, b) clusters that could contain more than 1 cell type (heterogeneous clusters indicated in Supplementary Table 3), c) clusters containing features of low quality cells, multiplets or central brain neurons (38, 85, 112, 102 and 120, see Supplementary Table 3), and d) clusters containing TE cells (220, 223, 224, and 233). We removed TE neurons because they are very different from other optic lobe neurons11, and might present unique regulatory relationships that would complicate model building. The removed heterogenous clusters are based on previous work11, and on results obtained by grouping our P50 optic lobe atlas11 with another optic lobe atlas14, produced at P48: this increase in cell number allowed splitting some of our previously published clusters. By “main OPC clusters”, we mean the subset of “All clusters” that were identified as being produced in the main OPC28 and for which we were able to assign a spatial, temporal and Notch origin (see Supplementary Table 1).

Production of the gene expression tables

We filtered the scRNA-seq datasets of the optic lobe produced previously11 at P15, P30, P40, P50, P70 and adult stage to keep only “All clusters” or “main OPC clusters” as detailed above, and to keep only the subset of genes that we considered as unambiguously expressed, i.e., with mRNA detected in at least 10% of the cells for at least 1 neuronal cluster. We then used the AverageExpression function of Seurat_4.3.0 to produce tables containing the log-normalized average gene expression of each cluster. We used averaged gene expression for our analyses rather than single-cell gene expression to mitigate the effect of dropouts (i.e., false negative gene expression due to randomness in the capture of mRNA). We also produced additional tables in which, for each gene, we shuffled the values between clusters to produce log-normalized randomized average gene expression.

Binarization of gene expression

Several steps of our analyzes require binarization of gene expression. To do so, we used our previously published mixture modeling tables, which quantify the probability of expression for each gene in each cluster of our scRNA-seq atlas11. We binarized these expression probabilities using a threshold of 0.5, and we refer to these binarized tables as “MM0.5” hereafter.

Selection of the predictors

To select the predictor genes, we further filtered the log-normalized average gene expression tables for TFs, CSSs or TSCs. For TFs, we used a list containing “TFs with characterized binding domains, computationally predicted (putative) TFs, chromatin-related proteins and transcriptional machinery components”50. For CSSs, we used a list of “Drosophila CSS proteins potentially involved in cell recognition” obtained by performing “BLAST searches with extracellular […] domain sequences from a variety of species and collated published information”51. Lastly, for TSCs we used the list we previously published, where TSCs were identified based on their sustained expression throughout development17. Moreover, for each list at each stage, we kept only the genes expressed at least once, and in fewer than 75% of either “All clusters” or “main OPC clusters”, according to MM0.5. Indeed, since our goal was to predict neuronal features only present in a subset of the clusters (differentially expressed genes, non-pan-neuronal morphologies) and to find candidate regulators for them, it made sense to focus on differentially expressed predictors. Moreover, pan-neuronal predictors will always be well correlated with any feature that is present is many neurons, even though they are not involved in its regulation. We produced the tables of predictors using either regular or randomized gene expression.

The developmental origins used as predictors are found in Supplementary Table 1 for temporal and Notch origins. However, some clusters can be produced in more than one spatial domain. For our analyzes, we thus used Supplementary Table 1 to assign a unique origin to each neuronal cluster among vDpp, vdDpp, dDpp, vDpp.vOptix, vdDpp.vdOptix, dDpp.dOptix, vOptix, vdOptix, dOptix, vOptix.vVsx, vdOptix.vdVsx, dOptix.dVsx, vVsx, vdVsx, dVsx, vVsx.dVsx.dOptix, vOptix.vVsx.dVsx and whole_OPC. For instance, a cluster produced in both the vDpp and the vOptix domains was assigned the origin “vDpp.vOptix” for our analyses. We also removed a few clusters, either because we determined they were not produced from the main OPC, or because we could not determine their spatial origin (11, 36, 21, 156, 177, 233, see Supplementary Table 1). We shuffled which clusters are produced in each spatial, temporal and Notch origin to produce randomized tables of developmental origins.

Importanly, although vDpp and dDpp contain clusters produced exclusively in these domains, because of the lack of dataset for neurons from the dDpp domain in most cases we could not identify neuronal types produced in the vDpp but not dDpp. Therefore, some neurons in vdDpp might be in fact produced exclusively in the vDpp domain. Such reasoning also applies to the vVsx, dVsx and vdVsx origins.

Production of the training and testing sets for molecular features

To choose which clusters and genes to use for random forest predictions, we further filtered the log-normalized average gene expression tables, in “All clusters” and “main OPC clusters independently. Both training and testing sets should contain examples of positive and negative regulation of the feature of interest, and therefore all genes used to produce a random forest model should be expressed in at least 2 clusters. In addition, genes expressed in only 1 cluster are perfectly correlated with the combination of TSCs specific to this cluster11,17, which are their best candidate regulators. Thus, we removed all genes expressed in 0 or 1 clusters according to MM0.5. We also removed all genes we considered pan-neuronal, i.e., expressed in more than 75% of the clusters according to MM0.5, because these could be regulated by any pan-neuronal TFs and a random forest model would therefore not be informative of their regulation. For each gene, we then produced a training set containing a randomly chosen subset of 80% of the clusters and a test set containing the remaining 20%, making sure that both the training and testing sets contained at least one cluster expressing the gene and one not expressing the gene, according to MM0.5.

Production of the training and testing sets for morphological features

For molecular features, we also treated independently “All clusters” and “main OPC clusters”. We chose only the features evaluated (either present or absent) in more than 10 neuronal clusters. Similarly to genes, we also selected only the morphological features present in at least 2 clusters. Moreover, we selected only qualitative features, because the models produced for quantitative features would have to be evaluated in a different way. For each morphological feature we then produced a training set containing a randomly chosen subset of 80% of the clusters and a test set containing the remaining 20%. We made sure that both the training and testing sets contained at least one cluster presenting the feature and one not presenting the feature. The test set therefore contain at least 2 clusters, and the train set at least 8 clusters.

Production of the models

For each stage and each feature (molecular or morphological), we used the expression of different predictors (TFs, CSSs, TSCs, TFs that are not TSCs, developmental origins), in “All clusters” or “main OPC clusters”, with regular or randomized expression, to train a random forest model in the corresponding training set. If the feature to predict was part of the list of predictors, the feature was first removed from the predictors. We used the function randomForest with default parameters except for the number of trees that we set at 1000, using regression to predict gene expression or classification to predict morphological features (therefore we only built models for the morphologies that were assessed qualitatively and not quantitively), using the R package randomForest 4.6-14.

Evaluation of the models

To evaluate the quality of the model, we then computed either Pearson correlation (for molecular features) or MCC (for all features) between predicted values and real values, that we call observed values. We used the MCC because it gives the same importance to all 4 categories of the confusion matrix (true and false positives and negatives), which is well suited for unbalanced datasets52 such as differentially expressed genes or morphological features. Indeed, since we are interested in features present only in a subset of neuronal types, other metrics such as Pearson correlation would be disproportionately affected by the ability of the model to predict the absence of the feature (i.e., the majority of the clusters).

However, MCC requires binarized predicted and observed values. For morphological features, both the observed and predicted values were already binarized in the case of qualitative features. For molecular features, we used MM0.5 as observed values, and we binarized the predicted values. The predicted values range between a maximal and a minimal value, different for each gene, with higher values in clusters where the gene is most likely to be expressed according to the model. However, it is unclear which of these values correspond to expression and non-expression according to the model. Therefore, to binarize the predictions, we scaled them between 0 and 1, and then binarized them using all thresholds between 0 and 1 with an increment of 0.01 (0, 0.01, 0.02, etc.): a threshold of 0 would correspond to a gene always expressed in the test set, a threshold of 1 to a gene never expressed, both impossible cases based on how we produced the test sets. If the model made meaningful predictions, and intermediate threshold will give better results. We therefore used the threshold giving the best fit between predicted and observed values, i.e., the one giving the highest possible MCC. MCC ranges between 1 (perfect correlation) and −1 (perfect anticorrelation), with 0 denoting lack of correlation. However, the MCC values we obtained are slightly inflated because we choose the best possible MCC for each feature: our results obtained with regular predictors should therefore be compared to the results obtained with randomized predictors, and not to an MCC of 0.

Result tables and their interpretation

For molecular features, for each stage we produced several tables corresponding to different combinations of type of predictors, type of expression (regular or random), and clusters used (“All” or “main OPC”). These table contain the following columns: 1) “Features”: the gene for which the model was built, 2) “Top_predictors”: a ranking of their best 30 predictors (less if there are less predictors), which are the most likely candidate regulators of this feature, 3) “%IncMSE”: the percent increase in mean square error of the model when the values of the predictor are shuffled 4) “RMSE_TestSet”: The square root of the mean squared error between observed and predicted values, 5) “Cor_TestPreds_TestObs”: Pearson correlation between observed and predicted values, 6) “Best_MCC_TestSetBin_TestSetMM”: best MCC obtained between observed and predicted values, and 7) “Best_MCC_Threshold”: the threshold giving the best MCC for this model. For morphological features, we produced similar tables, with the following columns: 1) “Features”, 2) “Top_predictors”, 3) “%IncMSE”, 4) “OOBE”: Out Of Bag estimate of error rate, 5) “MCC_TestObs_TestPreds”.

For any feature and each type of predictors (TFs, CSSs, developmental origin…), high MCC values indicate that the random forest algorithm was able to identify a correlation between the predictors and the feature that is valid both in the train and the test sets. Different train set would lead to different models, and different test set would lead to different MCC values for a given model. Moreover, random forest functions by building a forest of many decision trees built from a randomly sampled subset of the data: the algorithm will produce slightly different models, and therefore obtain a slightly different MCC value, each time it is run. By chance, sometimes this MCC value will be a particularly high or low outlier. Lastly, sometimes correlation between predictors and a feature can be due to chance and not be the result of a molecular mechanism, which explains why randomized predictors sometimes yield models with high MCC values. To mitigate these caveats, in this work for each feature we used a randomly chosen train and test set, and we only drew conclusions by using all the MCC values obtained for the hundreds of molecular and morphological features tested. Moreover, for molecular features, the correlations we used were obtained across dozens of cell types (all main OPC or all optic lobe neurons). Together, this avoids drawing conclusions based on results only valid in a specific set of cell types. However, because the status of many morphological features could be evaluated only in few clusters, in several cases the train and test set were particularly small (in the worst cases, 8 and 2 clusters respectively). This strongly increased the chance that predicted status of these molecular features were the same in each cluster of the test set, and in such cases MCC is not defined. Because of this, models could be produced for only a small number of morphological features, which limits the strength of the conclusions drawn from morphological features.

For anyone interested in using the results of our random forest approach to experimentally find regulators of a given feature, it should be noted that all features present in our result tables will have corresponding ranked predictors, but only those associated to a random forest algorithm with good MCC values should be considered as candidate regulators of the feature. As discussed in the main text, models with MCC of at least 0.745 can be considered as yielding high quality candidates. However, it is very likely that many models below this MCC also yield high quality candidate regulators for their target gene. Since random forest only identifies correlations between the presence of a predictors and of a target feature, even very good candidate regulators could be false positives. It would therefore be useful to use other approaches to reduce this number of false positives: for instance, if a TF is candidate regulator of a gene, this could be done by exploring the presence of binding sites for the TF near the promoter of this gene, or ChIP-seq data even if it was produced in other tissues or developmental stages.

Lastly, the models produced using only main OPC clusters performed, on average, better than models produced using all neuronal clusters (i.e., including main OPC clusters, but also tOPC, inner proliferation center, lamina, central brain… Supplementary Fig. 5), which suggests that the regulation of neuronal features is different in neurons produced from different neurogenic domains. Therefore, it would be advisable to use the models produced using all neuronal clusters only when necessary, i.e., to study regulations in non-main OPC clusters.

Heatmaps of the presence of neuronal features

Production of the heatmaps

For all the heatmaps, the list of TFs, CSSs or TSCs used where the same as for the random forest analyses (see section “Selection of the predictors”).

To produce heatmaps of the TSCs expressed during all of development in all clusters sharing a combination of specification modules, we only used the TSCs expressed in less than 50% of the main OPC neuronal clusters according to MM0.5. We considered that a TSC was expressed in a given cluster during all its development if it was expressed in at least 5 out of the 6 developmental stages of our scRNA-seq atlas according to MM0.5. We did not require for 6 out of 6 stages to account for imperfection in the modeling of gene expression by Mixture Modeling.

To produce heatmaps of genes expressed at a given stage in all clusters sharing a combination of specification modules, we only used genes expressed in less than 50% of the main OPC neuronal clusters according to MM0.5.

To produce heatmaps of morphological features present in all clusters sharing a combination of specification modules, if the presence of a feature in a neuronal type was unsure, we considered it absent (i.e., all “NA” values were replaced by a value of 0). Indeed, we chose to be conservative because the purpose of these heatmaps is to find features present in all neuronal types sharing a common origin.

In any case, we considered that a feature (gene, TSCs, morphological feature) was present in all clusters from a given origin when 75% of the clusters from this origin presented the feature, instead of 100%, to account for the imperfect resolution of our determination of developmental origins.

Proportion of the TSCs used to produce the heatmaps

There are 65 TSCs that are expressed in at least one main OPC neuronal cluster for at least 5 out of the 6 developmental stages of our scRNA-seq atlas according to MM0.5. Out of these, 12 are continuously expressed in a single cluster: foxo, sr, fru, Camta, dve, CG43689, run, CG34340, ab, bsh, tsh, and zfh2. These TSCs were not taken into account for the production of the heatmaps because it is not possible to compare their expression across several neurons from the same origin: they are either expressed in less than 75% of any group of more than 2 neurons sharing a specification module, or expressed in a neuron with a unique combination of temporal, spatial and Notch origin (e.g., bsh is expressed in Mi1, the only Hth NotchON neuron produced from the whole main OPC). This leaves 53 TSCs continuously expressed in more than one cluster for which we could look for a correlation between their expression and the developmental origin of neurons. Out of these, 35 (66%) were expressed in at least one group of neurons sharing the same combination of specification modules, and 18 (34%) were not (Extended Data Fig. 9).

Using the heatmaps to infer candidate regulatory relationships.

The purpose of these heatmaps is to highlight gene expression or morphological features that are present in all neurons from one or several developmental origins. This is figured in bright blue. It can suggest regulatory relationships: if all neurons sharing a developmental origin express a TSC, the expression of the TSC could be activated by the specification factors characterizing this developmental origin (e.g., NotchON activates ap expression, see main text). Similarly, if all these neurons also express a given gene, the expression of the gene could be activated by a TSC present in all neurons from this origin (e.g., NotchON activates ap expression, and ap activates a cholinergic phenotype in all temporal windows except the last one, see main text).

For each neuronal feature, however, it is important to consider the largest group of neurons that share both the presence of the feature and a given combination of specification modules. For instance, the TSC ap is expressed in all the six dOptix clusters (Extended Data Fig. 7a): one could therefore hypothesize that ap is turned on by the dOptix specification module. However, all thirty-nine NotchON clusters, regardless of spatial origin, express ap (Fig. 4c). This indicates that ap expression is not controlled by spatial patterning and is instead controlled by Notch signaling, as was previously shown25. These heatmaps also highlight the interdependency between TSC expression and the combinations of specification modules that produce neuronal types. For instance, all dOptix-only clusters are either from the Hbn/Opa/Slp or the Slp/D temporal windows (Extended Data Fig. 8b). Neurons from other temporal windows are therefore never produced exclusively from the dOptix domain: they must either be produced from larger regions that encompass dOptix (for instance Tm3 is from the Erm/Ey temporal window and is produced across the Optix and Vsx domains, Fig. 2f), or be produced exclusively from the dOptix domains but be culled. This also implies that no TSC is under the control of dOptix associated with temporal windows other than Hbn/Opa/Slp or Slp/D (since no neuronal type is produced from these combinations of specification modules).

Lastly, even if a feature is present in all neurons from one or several developmental origins, it can also be present in a subset of the neurons from other developmental origins. This is represented in medium blue, and can be interpreted in two different ways. First, when only a subset of neurons sharing a specification module express a feature, they might share exclusively another specification module. This can be illustrated with the TSC TfAP-2. It is expressed in only 4/8 neurons of the Hth/Opa temporal windows, and they are thus represented in medium blue on Extended Data Fig. 7b (where neurons are grouped by temporal windows). However, these 4 neurons are the only NotchON neurons of the Hth/Opa temporal window, and they are thus represented in light blue on Extended Data Fig. 8a (where neurons are grouped by temporal and Notch origin). However, sometimes when only a subset of the neurons from a developmental origin presents a feature, it is not possible to find a specification module common to these neurons. In that case, our heatmaps do not provide candidate regulatory mechanism for expression in these neurons.

Statistics & Reproducibility

No statistical method was used to predetermine sample size, the number of cells sequenced from each main OPC domain was chosen based on the relative size of these different domains. Individual Drosophila flies were chosen randomly for all experiments. We did not exclude data, except some single-cell transcriptomes according to the quality control steps described in Methods. The investigators were not blinded to allocation during experiments and outcome assessment. All immunostainings were performed on at least 3 biological replicates (brain of different flies).

Extended Data

Extended Data Fig. 1: Characterization of the lines used to identify the spatial origin of main OPC neurons by scRNA-seq.

Extended Data Fig. 1:

A) Schematic describing how a temperature-sensitive (ts) memory cassette (MC) can be used to label all progeny of the cells expressing Gal4 (here driven by Optix-Gal4) while they are placed at the permissive temperature of 29°C but not at the restrictive temperature of 18°C. Placing the flies at 29°C only during neuronal specification prevents later activation of the memory cassettes in postmitotic neurons (see Extended Data Fig. 1b). Due to technical difficulties, some of our memory cassette lacked the tub-Gal80ts construct: they labeled any cell that ever expressed Gal4, at any temperature, as well as its progeny.

B) Hoechst labelling of an adult optic lobe, and expression pattern of Optix-ts-MC and hh-ts-MC, in adult flies grown exclusively at 18°C. All images are maximum intensity z-projections of 15 slices spanning the whole optic lobe, showing that the number of cells labelled in Optix-ts-MC and hh-ts-MC is negligible compared to the number of optic lobe nuclei, and that the Gal80ts efficiently prevents activation of the MC. D: dorsal, L: lateral, M: medial, V: ventral. Dashed lines: optic lobe, with the lamina excluded.

C) Top: Optix-ts-MC third-instar larval expression pattern showing that some neurons from the Optix domain express Vsx1. Middle: additional view from the same brain showing that the Vsx domain neuroepithelial cells (high Vsx1+) are not labelled by the Optix-ts-MC line. Bottom: the vOptix-MC line is ectopically expressed in neuroepithelial cells from the vVsx stripe (circled in orange), which we accounted for in later analyses (Methods). N: neurons, NE: neuroepithelium. White dashed line: GFP expressed in the main OPC. For orientation, see schematics and Fig. 1c.

D) In about 25% of the optic lobes, vOptix-MC is ectopically expressed in the ventral tip of the OPC (which is not part of the Optix domain), which we accounted for in later analyses (Methods). Labelled neurons outside of the dashed lines are not part of the optic lobe. For orientation, see schematic and Fig. 1c.

E) In some optic lobes, pxb-MC labels parts of the main OPC that were not produced in the Vsx domain (this shows an extreme example, at early pupal stage). We eliminated these cases by visually selecting optic lobes that did not exhibit ectopic expression before FACS and sequencing (Methods). For orientation, see schematic and Fig. 1c.

F) pxb-MC expression pattern, showing that pxb-MC labels a band of Vsx1+ cells that corresponds to the Vsx domain at early pupal stage, and that pxb-MC does not label the Vsx1+ cells scattered in the Optix domain that can also be seen on Extended Data Fig. 1c. For orientation, see schematic and Fig. 1c.

(C-F) The immunostainings were done following the protocol summarized on Fig. 2b. Unless indicated, all scale bars are 20 μm.

Extended Data Fig. 2: Confidence in cell annotations and calculation of a Class enrichment score in the scRNA-seq libraries.

Extended Data Fig. 2:

A) Confidence in the annotation of each cell, grouped by Class, of one of the Optix libraries. Each dot is a cell, and the violin plots show their distribution. This shows that although annotation confidence is high for most Classes, it is sometimes lower. This can have biological reasons, and it is for instance expected for Classes corresponding to transcriptionally similar neurons (e.g., Dm3a/b and Tm9v/d that are highlighted on the figure) since the neural network uses transcriptional information to distinguish between cell types. Classes with more than 80% of the cells with a score below 0.5 are prefixed with “LowC”, i.e., “Low Confidence”.

B) Schematic explaining the rationale of abundance normalization and enrichment scores. The normalized abundance of a neuron exclusively produced in the Vsx domain is higher in a dataset produced from the Vsx domain, i.e., in the pxb dataset (even though the number of cells, here 8 pink cells, is the same across both datasets). We used our previously published scRNA-seq atlas as a reference dataset, since it represents a largely unbiased sampling of optic lobe neurons from all main OPC domains, which represents a largely unbiased sampling of optic lobe neurons from all main OPC domains.

C) Normalized abundance of the indicated Medulla intrinsic (Mi) neurons in the different libraries.

Extended Data Fig. 3: Experimental validation of the spatial origin of main OPC neurons.

Extended Data Fig. 3:

This figure shows the domain of expression of various TFs, used as markers to confirm spatial origins obtained from scRNA-seq data. The markers used for each cell type, and how their expression was used to assess spatial origin, is detailed in Table S1.

A) Expression of the indicated TFs (y-axis) in the indicated neuronal clusters (x-axis), at P15. For each neuronal cluster, their markers can be expressed in a main OPC domain larger than their spatial origin, because most of the markers are expressed in several neuronal clusters (some not shown in this panel).

B-J) Expression pattern of the indicated markers in the main OPC at the L3 stage. Grey dashed line: outline of the optic lobe, white dashed lines: localization of the labelled neurons (see Table S1), unless otherwise indicated. Scale bars are 20 μm.

D) Dashed line: estimation of the ventral border of the Vsx domain.

F) Cells in the Vsx domain are circled in green when they co-express the indicated markers, and circled in red when they do not. This shows that no cell expresses ey, kn and D in the Vsx domain, except maybe cells indicated by an arrowhead (these are uncertain because they are located at the limit of the Vsx domain). The white dotted lines correspond to the insets placed at the bottom, which are rotated 90 degrees clockwise to the original image.

I) The dashed white line indicates cells in the Vsx domain, while the solid white line indicates cells that might not be part of the Vsx domain.

J) The arrowhead indicates Dll+ Vsx1+ cells in the Vsx, dOptix and what is likely the dDpp domains.

Extended Data Fig. 4: Thresholds used to binarize the enrichment scores of main OPC neurons.

Extended Data Fig. 4:

A-F) Log-enrichment scores of the indicated Classes (x-axis) in each dataset. Enrichment scores were calculated by dividing the normalized abundances in the indicated datasets, by the normalized abundances in the Reference dataset. Vertical dotted lines: enrichment of the Class used as a threshold (indicated by horizontal dotted lines) to binarize the enrichment scores. Any class above this dotted line was considered as produced from the main OPC spatial domain corresponding to the dataset. Blue horizontal shaded regions: enrichment values close to the binarization threshold (Methods). The name of each neuronal Class is colored according to our validation experiments (Table S1). Red: neuronal types that should not be produced from the main OPC spatial domain corresponding to the indicated dataset, Green: neuronal types that could be produced from the main OPC spatial domain corresponding to the indicated dataset, Blue: neuronal types for which the presence in the corresponding main OPC domains was unclear (Table S1).

Extended Data Fig. 5: Binarized enrichment scores of the Classes in each dataset, and corresponding quality control metrics.

Extended Data Fig. 5:

A-B’) Green squares: Classes (y-axis) considered as enriched in the corresponding dataset (x-axis) according to the binarization thresholds indicated in Extended Data Fig. 4 (adult reference dataset) or Annex 1 (P15 reference dataset). P15 was used as reference dataset in addition to the adult stage because TE neurons die before adulthood. Grey squares: Classes considered as depleted in the corresponding dataset. A = at least one of the libraries of the dataset had low abundance, i.e., contained fewer than 3 cells of this Class, C = at least 80% of the cells of the Class were annotated with a confidence score strictly below 0.5, T = the enrichment of the Class was close to the binarization threshold (horizontal blue shaded region in Extended Data Fig. 4). Abbreviations used as Class names are explained in Table S3 (the Class names ending with a capital G are glial cells).

Extended Data Fig. 6: UMAP visualizations of the libraries obtained at early pupal stages.

Extended Data Fig. 6:

A-C) UMAP plots obtained using 150 principal components. The cells are labeled and colored according to the key presented on the bottom right. C) Red arrow highlights Dm11 (see Table S1).

Extended Data Fig. 7: Terminal selector candidates continuously expressed in all neurons sharing a specification module.

Extended Data Fig. 7:

A-C) Terminal selector candidates (y-axis) continuously expressed in more than 75% of main OPC clusters (x-axis) from at least one spatial, temporal or Notch specification module. We used a threshold of 75% of clusters expressing continuously a TSC instead of 100% to account for the imperfect resolution of our determination of developmental origins. Light blue: Expression sustained during development in this cluster and in at least 75% of clusters from the same combination of specification modules, Medium blue: Expression sustained during development in this cluster and in less than 75% of clusters from the same combination of specification modules, Dark blue: gene not expressed throughout development (See Methods on how to interpret these heatmaps).

Extended Data Fig. 8: Terminal selector candidates continuously expressed in all neurons sharing a combination of two specification modules.

Extended Data Fig. 8:

A-C) Terminal selector candidates (y-axis) continuously expressed in more than 75% of main OPC clusters (x-axis) from at least one combination of two specification modules. We used a threshold of 75% of clusters expressing continuously a TSC instead of 100% to account for the imperfect resolution of our determination of developmental origins. Combinations of specification modules corresponding to a single cluster were not plotted. Colors as in Extended Data Fig. 7 (See Methods on how to interpret these heatmaps).

Extended Data Fig. 9: Terminal selector candidates continuously expressed in neurons sharing a combination of specification modules.

Extended Data Fig. 9:

A) Terminal selector candidates (y-axis) continuously expressed in more than 75% of main OPC clusters (x-axis) from at least one combination of spatial, temporal and Notch origins. We used a threshold of 75% of clusters expressing continuously a TSC instead of 100% to account for the imperfect resolution of our determination of developmental origins. Combinations of specification modules corresponding to a single cluster were not plotted. Colors as in Extended Data Fig. 7 (See Methods on how to interpret these heatmaps).

B) Number of clusters expressing continuously each of the 53 TSCs present in more than one cluster of the main OPC. Out of these, 66% (in pink) were expressed in at least one group of neurons sharing the same combination of specification modules (Methods). The TSCs that were not associated with any combination of specification modules (in blue) are expressed in a much lower number of clusters (median of 3 compared to 16). Therefore, only small groups of neurons can share both their expression and a developmental origin. Characterizing more precisely the developmental origin of neurons, for instance by splitting temporal windows or spatial domains, would therefore likely allow finding groups of neurons sharing the same combination of specification modules and the expression of one of the remaining (blue) TSCs.

Extended Data Fig. 10: Regulation of terminal selector candidates expression by the specification factors.

Extended Data Fig. 10:

When all progenitors expressing a given specification factor (ex: erm for the first row) give rise to neuronal types that all express the same TSC (ex: kn for the first row), this specification factor might be one of the activators of the expression of this TSC. In that case, preventing the expression of this specification factor in progenitors should prevent the expression of this TSC in neurons (and the reciprocal, for repressive interactions). Here, we present occurrences where this was experimentally validated. Throughout the figure, we interpreted the lack of labeling for a given TSC as a loss of expression, although it could also be due to the death of the neurons usually expressing this TSC (in most cases, this is unlikely since the use of mutant clones or RNAi clones shows that neurons are still produced when the specification factor is not expressed or functional). All the graphs are extracted from Extended Data Fig. 8, in which neuronal clusters are grouped according to temporal window and Notch signaling, and we also indicated when neuronal clusters expressing a given terminal selector candidate are present on Extended Data Fig. 8 but not on this figure. However, the combination of a temporal and a Notch origin that produce a single neuronal cluster were not plotted on Extended Data Fig. 8 nor on this figure. Because spatial patterning factors are likely to regulate terminal selector expression through changes in chromatin accessibility, the mechanism of this regulation is more difficult to study and we did not find already published experimental validations for our candidates.

Supplementary Material

Supplementary Figure 1

Supplementary Fig. 1: Adult morphological features present in all neurons sharing various combinations of specification modules.

Adult morphological features (y-axis) present in more than 75% of main OPC clusters (x-axis) from at least one of the following: Temporal (A), Spatial (B) or Notch (C) origin, or from at least one combination of all three specification modules (D). We used a threshold of 75% of clusters presenting a morphological feature instead of 100% to account for the imperfect resolution of our determination of developmental origins. The names of the morphological features are explained in Table S2. Combinations of specification modules corresponding to a single cluster were not plotted (because in this case it is impossible to compare morphology across several neurons from the same origin). Light blue: Present in this cluster and in at least 75% of clusters from the same combination of specification modules, Medium blue: Present in this cluster and in less than 75% of clusters from the same combination of specification modules. Dark blue: lack of the morphological feature or presence unsure. Some neuronal types could not be included in these analyzes because their corresponding neuronal cluster is unknown: these results could change when more clusters are annotated. (See Methods on how to interpret these heatmaps). The definition of each feature is indicated in the column “Comments” and the line corresponding to the feature in Table S2.

Supplementary Figure 2

Supplementary Fig. 2: Adult morphological features present in all neurons sharing combinations of two specification modules.

Adult morphological features (y-axis) present in more than 75% of main OPC clusters (x-axis) from at least one combination of two developmental origins. We used a threshold of 75% of clusters presenting a morphological feature instead of 100% to account for the imperfect resolution of our determination of developmental origins. The names of the morphological features are explained in Table S2. Combinations of specification modules corresponding to a single cluster were not plotted (because in this case it is impossible to compare morphology across several neurons from the same origin). Colors as in Extended Data Fig. 11. Some neuronal types could not be included in these analyzes because their corresponding neuronal cluster is unknown: these results could change when more clusters are annotated. (See Methods on how to interpret these heatmaps). The definition of each feature is indicated in the column “Comments” and the line corresponding to the feature in Table S2.

Supplementary Figure 3

Supplementary Fig. 3: Specification modules and terminal selector candidates encode similar information.

A-F) Matthews Correlation Coefficient between observed and predicted presence of a neuronal feature. The type of neuronal feature predicted and the type of predictor used are indicated on each plot. The specific TFs are all the TFs differentially expressed between neuronal types. See Fig. 6 for boxplots legend. NO = Notch Origin, SO = Spatial Origin, TO = Temporal Origin.

B) Dark blue line: median of the MCC obtained when using specific TF expression as predictors, Light blue line: median of the MCC obtained when using randomized specific TFs expression as predictors, Green line: median of the MCC obtained when using developmental origin (SO+TO+NO) as predictors.

C-D) Some neuronal types could not be included in these analyzes because their corresponding neuronal cluster is unknown: these results could change when more clusters are annotated.

Supplementary Figure 4

Supplementary Fig. 4: Specification modules may not be remembered epigenetically after TSC combinations are established.

A, B) Matthews Correlation Coefficient between observed and predicted presence of a neuronal feature. The type of neuronal feature predicted and the type of predictor used are indicated on each plot. The specific TFs are all the TFs differentially expressed between neuronal types. See Fig. 6 for boxplots legend. SO = Spatial Origin, TO = Temporal Origin, NO = Notch Origin. Dark blue line: median of the MCC obtained when using specific TFs expression as predictors, Light blue line: median of the MCC obtained when using randomized specific TFs expression as predictors.

Supplementary Figure 5

Supplementary Fig. 5: Prediction of neuronal features using TF expression or CSS expression, in main OPC or in all neuronal clusters.

A-C) Matthews Correlation Coefficient between observed and predicted presence of a neuronal feature. The type of neuronal feature predicted and the type of predictor used are indicated on each plot. The specific TFs are all the TFs differentially expressed between neuronal types. See Fig. 6 for boxplots legend.

B-C) Some neuronal types could not be included in these analyzes because their corresponding neuronal cluster is unknown: these results could change when more clusters are annotated.

Supplmentary table 1
Supplementary Table 3
Supplementary discussion
Supplementary table 2
Supplementary Data 2
Supplementary Code 2
Supplementary Data1
Supplementary Code 1
Supplementary Data 3

Acknowledgments

We thank all members of the Desplan and Konstantinides laboratories for helpful feedback while working on the manuscript. The Desplan laboratory was supported by grants from the National Eye Institute R01EY017916 and R01EY13010, and by Tamkeen under the NYU Abu Dhabi Research Institute Award to the NYUAD Center for Genomics and Systems Biology (ADHPG-CGSB). Research in the Konstantinides lab is supported by the European Research Council (ERC) under the European Union’s Horizon 2020 research and innovation program (grant agreement No. 949500). F.S. was supported by New York University (MacCracken Fellowship and Dean’s Dissertation Ph.D. Fellowship). I.H. was supported by a senior postdoctoral fellowship from the Kimmel Center for Stem Cell Biology and by a Marie Skłodowska-Curie Postdoctoral Fellowship (101154260). Y.-C.C. was supported by New York University (MacCracken Fellowship), by a NYSTEM institutional training grant (Contract #C322560GG), and by a Scholarship to Study Abroad from the Ministry of Education, Taiwan. J.M. was supported by NIH fellowships from the National Eye Institute and the BRAIN initiative (F32EY028012 and K99EY032269, respectively). P.V. was supported the Vision Science Research Program (University of Toronto: Ophthalmology and Vision Sciences and UHN), the Ontario Graduate Scholarship, and the Queen Elizabeth II/Pfizer Graduate Scholarship in Science and Technology. T.E. was supported by an NSERC Discovery Grant (RGPIN2015-06457). M.N.O. has been supported by K99/R00NS125117 from the National Institute of Neurological Disorders and Stroke, and Stowers Institute for Medical Research. This work was supported in part through the NYU IT High Performance Computing resources, services, and staff expertise. The funders had no role in study design, data collection and analysis, decision to publish or preparation of the manuscript.

Footnotes

Competing Interests Statement:

The authors declare no competing interests.

Data availability

All raw and processed data are available in GEO (GSE254562).

Code availability

The scripts used to process and visualize the data are available in Supplementary Code 1, 2.

References

  • 1.Fischbach K-F & Dittrich APM. The optic lobe of Drosophila melanogaster. I. A Golgi analysis of wild-type structure. Cell Tissue Res. 258, 441–475 (1989). [Google Scholar]
  • 2.Wu M et al. Visual projection neurons in the Drosophila lobula link feature detection to distinct behavioral programs. eLife 5, e21022 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3.Nern A, Pfeiffer BD & Rubin GM. Optimized tools for multicolor stochastic labeling reveal diverse stereotyped cell arrangements in the fly visual system. Proc. Natl. Acad. Sci. 112, E2967–E2976 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.Otsuna H & Ito K. Systematic analysis of the visual projection neurons of Drosophila melanogaster. I. Lobula-specific pathways. J. Comp. Neurol. 497, 928–958 (2006). [DOI] [PubMed] [Google Scholar]
  • 5.Rivera-Alba M et al. Wiring economy and volume exclusion determine neuronal placement in the Drosophila brain. Curr. Biol. CB 21, 2000–2005 (2011). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Takemura S et al. Synaptic circuits and their variations within different columns in the visual system of Drosophila. Proc. Natl. Acad. Sci. U. S. A. 112, 13711–13716 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.Takemura S et al. A visual motion detection circuit suggested by Drosophila connectomics. Nature 500, 175–181 (2013). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.Dorkenwald S et al. Neuronal wiring diagram of an adult brain. Nature 634, 124–138 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.Matsliah A et al. Neuronal parts list and wiring diagram for a visual system. Nature 634, 166–180 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.Nern A et al. Connectome-driven neural inventory of a complete visual system. Nature 1–13 (2025) doi: 10.1038/s41586-025-08746-0. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11.Özel MN et al. Neuronal diversity and convergence in a visual system developmental atlas. Nature 589, 88–95 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Konstantinides N et al. Phenotypic Convergence: Distinct Transcription Factors Regulate Common Terminal Features. Cell 174, 622–635.e13 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.Davis FP et al. A genetic, genomic, and computational resource for exploring neural circuit function. eLife 9, e50901 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.Kurmangaliyev YZ, Yoo J, Valdes-Aleman J, Sanfilippo P & Zipursky SL. Transcriptional Programs of Circuit Assembly in the Drosophila Visual System. Neuron S0896627320307741 (2020) doi: 10.1016/j.neuron.2020.10.006. [DOI] [PubMed] [Google Scholar]
  • 15.Simon F & Konstantinides N. Single-cell transcriptomics in the Drosophila visual system: Advances and perspectives on cell identity regulation, connectivity, and neuronal diversity evolution. Dev. Biol. 479, 107–122 (2021). [DOI] [PubMed] [Google Scholar]
  • 16.Ferreira AAG & Desplan C. An Atlas of the Developing Drosophila Visual System Glia and Subcellular mRNA Localization of Transcripts in Single Cells. 2023.08.06.552169 Preprint at 10.1101/2023.08.06.552169 (2023). [DOI] [Google Scholar]
  • 17.Özel MN et al. Coordinated control of neuronal differentiation and wiring by sustained transcription factors. Science 378, eadd1884 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Arain U, Islam IM, Valentino P & Erclik T. Concurrent temporal patterning of neural stem cells in the fly visual system. 2022.10.13.512100 Preprint at 10.1101/2022.10.13.512100 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.Chen Y-CD et al. Using single-cell RNA sequencing to generate predictive cell-type-specific split-GAL4 reagents throughout development. Proc. Natl. Acad. Sci. 120, e2307451120 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.Malin JA, Chen Y-C, Simon F, Keefer E & Desplan C. Spatial patterning controls neuron numbers in the Drosophila visual system. Dev. Cell 59, 1132–1145.e6 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.El-Danaf RN et al. Morphological and functional convergence of visual projections neurons from diverse neurogenic origins in Drosophila. 2024.04.01.587522 Preprint at 10.1101/2024.04.01.587522 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Sen SQ. Generating neural diversity through spatial and temporal patterning. Semin. Cell Dev. Biol. 142, 54–66 (2023). [DOI] [PubMed] [Google Scholar]
  • 23.Holguera I & Desplan C. Neuronal specification in space and time. Science 362, 176–180 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24.Valentino P & Erclik T. Spalt and disco define the dorsal-ventral neuroepithelial compartments of the developing Drosophila medulla. Genetics 222, iyac145 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25.Erclik T et al. Integration of temporal and spatial patterning generates neural diversity. Nature 541, 365–370 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26.Li X et al. Temporal patterning of Drosophila medulla neuroblasts controls neural fates. Nature 498, 456–462 (2013). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27.Suzuki T, Kaido M, Takayama R & Sato M. A temporal mechanism that produces neuronal diversity in the Drosophila visual center. Dev. Biol. 380, 12–24 (2013). [DOI] [PubMed] [Google Scholar]
  • 28.Konstantinides N et al. A complete temporal transcription factor series in the fly visual system. Nature 604, 316–322 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29.Zhu H, Zhao SD, Ray A, Zhang Y & Li X. A comprehensive temporal patterning gene network in Drosophila medulla neuroblasts revealed by single-cell RNA sequencing. Nat. Commun. 13, 1247 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30.Chang T, Mazotta J, Dumstrei K, Dumitrescu A & Hartenstein V. Dpp and Hh signaling in the Drosophila embryonic eye field. Development 128, 4691–4704 (2001). [DOI] [PubMed] [Google Scholar]
  • 31.Hasegawa E et al. Concentric zones, cell migration and neuronal circuits in the Drosophila visual center. Development 138, 983–993 (2011). [DOI] [PubMed] [Google Scholar]
  • 32.Hobert O & Kratsios P. Neuronal identity control by terminal selectors in worms, flies, and chordates. Curr. Opin. Neurobiol. 56, 97–105 (2019). [DOI] [PubMed] [Google Scholar]
  • 33.Hobert O. Homeobox genes and the specification of neuronal identity. Nat. Rev. Neurosci. 22, 627–636 (2021). [DOI] [PubMed] [Google Scholar]
  • 34.Murgan S et al. Atypical Transcriptional Activation by TCF via a Zic Transcription Factor in C. elegans Neuronal Precursors. Dev. Cell 33, 737–745 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 35.Filippopoulou K, Couillault C & Bertrand V. Multiple neural bHLHs ensure the precision of a neuronal specification event in Caenorhabditis elegans. Biol. Open 10, bio058976 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 36.Bertrand V & Hobert O. Linking asymmetric cell division to the terminal differentiation program of postmitotic neurons in C. elegans. Dev. Cell 16, 563–575 (2009). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 37.Deneris ES & Wyler SC. Serotonergic transcriptional networks and potential importance to mental health. Nat. Neurosci. 15, 519–527 (2012). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 38.Zhang XL, Spencer WC, Tabuchi N, Kitt MM & Deneris ES. Reorganization of postmitotic neuronal chromatin accessibility for maturation of serotonergic identity. eLife 11, e75970 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 39.Maximilien C & Claude D. Coordination between stochastic and deterministic specification in the Drosophila visual system. Science 366, (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 40.Naidu VG et al. Temporal progression of Drosophila medulla neuroblasts generates the transcription factor combination to control T1 neuron morphogenesis. Dev. Biol. 464, 35–44 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 41.Zhang Y, Lowe S, Ding AZ & Li X. Axon targeting of Drosophila medulla projection neurons requires diffusible Netrin and is coordinated with neuroblast temporal patterning. Cell Rep. 42, 112144 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 42.Court R et al. Virtual Fly Brain—An interactive atlas of the Drosophila nervous system. Front. Physiol. 14, (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 43.Holguera I et al. Temporal and Notch identity determine layer targeting and synapse location of medulla neurons. 2025.01.06.631439 Preprint at 10.1101/2025.01.06.631439 (2025). [DOI] [Google Scholar]
  • 44.Liaw A & Wiener M. Classification and Regression by randomForest. R News 2, 18–22 (2002). [Google Scholar]
  • 45.Chen Y-C & Konstantinides N. Integration of Spatial and Temporal Patterning in the Invertebrate and Vertebrate Nervous System. Front. Neurosci. 16, 854422 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 46.Xu C, Ramos TB, Rogers EM, Reiser MB & Doe CQ. Homeodomain proteins hierarchically specify neuronal diversity and synaptic connectivity. eLife 12, RP90133. [DOI] [PMC free article] [PubMed] [Google Scholar]

Methods-only references

  • 47.Gunawan F, Arandjelovic M & Godt D. The Maf factor Traffic jam both enables and inhibits collective cell migration in Drosophila oogenesis. Development 140, 2808–2817 (2013). [DOI] [PubMed] [Google Scholar]
  • 48.Chyb S & Gompel N. Wild-type morphology. in Atlas of Drosophila Morphology (eds Chyb S & Gompel N) 1–23 (Academic Press, San Diego, 2013). doi: 10.1016/B978-0-12-384688-4.00001-8. [DOI] [Google Scholar]
  • 49.Takemura S-Y, Lu Z & Meinertzhagen IA. Synaptic circuits of the Drosophila optic lobe: the input terminals to the medulla. J. Comp. Neurol. 509, 493–513 (2008). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 50.Rhee DY et al. Transcription Factor Networks in Drosophila melanogaster. Cell Rep. 8, 2031–2043 (2014). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 51.Kurusu M et al. A screen of cell-surface molecules identifies leucine-rich repeat proteins as key mediators of synaptic target selection in the Drosophila neuromuscular system. Neuron 59, 972–985 (2008). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 52.Chicco D & Jurman G. The advantages of the Matthews correlation coefficient (MCC) over F1 score and accuracy in binary classification evaluation. BMC Genomics 21, 6 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Supplementary Figure 1

Supplementary Fig. 1: Adult morphological features present in all neurons sharing various combinations of specification modules.

Adult morphological features (y-axis) present in more than 75% of main OPC clusters (x-axis) from at least one of the following: Temporal (A), Spatial (B) or Notch (C) origin, or from at least one combination of all three specification modules (D). We used a threshold of 75% of clusters presenting a morphological feature instead of 100% to account for the imperfect resolution of our determination of developmental origins. The names of the morphological features are explained in Table S2. Combinations of specification modules corresponding to a single cluster were not plotted (because in this case it is impossible to compare morphology across several neurons from the same origin). Light blue: Present in this cluster and in at least 75% of clusters from the same combination of specification modules, Medium blue: Present in this cluster and in less than 75% of clusters from the same combination of specification modules. Dark blue: lack of the morphological feature or presence unsure. Some neuronal types could not be included in these analyzes because their corresponding neuronal cluster is unknown: these results could change when more clusters are annotated. (See Methods on how to interpret these heatmaps). The definition of each feature is indicated in the column “Comments” and the line corresponding to the feature in Table S2.

Supplementary Figure 2

Supplementary Fig. 2: Adult morphological features present in all neurons sharing combinations of two specification modules.

Adult morphological features (y-axis) present in more than 75% of main OPC clusters (x-axis) from at least one combination of two developmental origins. We used a threshold of 75% of clusters presenting a morphological feature instead of 100% to account for the imperfect resolution of our determination of developmental origins. The names of the morphological features are explained in Table S2. Combinations of specification modules corresponding to a single cluster were not plotted (because in this case it is impossible to compare morphology across several neurons from the same origin). Colors as in Extended Data Fig. 11. Some neuronal types could not be included in these analyzes because their corresponding neuronal cluster is unknown: these results could change when more clusters are annotated. (See Methods on how to interpret these heatmaps). The definition of each feature is indicated in the column “Comments” and the line corresponding to the feature in Table S2.

Supplementary Figure 3

Supplementary Fig. 3: Specification modules and terminal selector candidates encode similar information.

A-F) Matthews Correlation Coefficient between observed and predicted presence of a neuronal feature. The type of neuronal feature predicted and the type of predictor used are indicated on each plot. The specific TFs are all the TFs differentially expressed between neuronal types. See Fig. 6 for boxplots legend. NO = Notch Origin, SO = Spatial Origin, TO = Temporal Origin.

B) Dark blue line: median of the MCC obtained when using specific TF expression as predictors, Light blue line: median of the MCC obtained when using randomized specific TFs expression as predictors, Green line: median of the MCC obtained when using developmental origin (SO+TO+NO) as predictors.

C-D) Some neuronal types could not be included in these analyzes because their corresponding neuronal cluster is unknown: these results could change when more clusters are annotated.

Supplementary Figure 4

Supplementary Fig. 4: Specification modules may not be remembered epigenetically after TSC combinations are established.

A, B) Matthews Correlation Coefficient between observed and predicted presence of a neuronal feature. The type of neuronal feature predicted and the type of predictor used are indicated on each plot. The specific TFs are all the TFs differentially expressed between neuronal types. See Fig. 6 for boxplots legend. SO = Spatial Origin, TO = Temporal Origin, NO = Notch Origin. Dark blue line: median of the MCC obtained when using specific TFs expression as predictors, Light blue line: median of the MCC obtained when using randomized specific TFs expression as predictors.

Supplementary Figure 5

Supplementary Fig. 5: Prediction of neuronal features using TF expression or CSS expression, in main OPC or in all neuronal clusters.

A-C) Matthews Correlation Coefficient between observed and predicted presence of a neuronal feature. The type of neuronal feature predicted and the type of predictor used are indicated on each plot. The specific TFs are all the TFs differentially expressed between neuronal types. See Fig. 6 for boxplots legend.

B-C) Some neuronal types could not be included in these analyzes because their corresponding neuronal cluster is unknown: these results could change when more clusters are annotated.

Supplmentary table 1
Supplementary Table 3
Supplementary discussion
Supplementary table 2
Supplementary Data 2
Supplementary Code 2
Supplementary Data1
Supplementary Code 1
Supplementary Data 3

Data Availability Statement

All raw and processed data are available in GEO (GSE254562).

The scripts used to process and visualize the data are available in Supplementary Code 1, 2.

RESOURCES