Skip to main content
iScience logoLink to iScience
. 2025 Mar 22;28(4):112276. doi: 10.1016/j.isci.2025.112276

Inverse design of a valley-Hall photonic topological insulator based on tandem residual neural networks

Bing-Jiang Wang 1, Le Zhang 1,5,∗, Ben-Xin Wang 2,∗∗, Dong-Ping Zhang 1, Ya-Guang Xie 3, Jin-Hui Cai 4
PMCID: PMC12005334  PMID: 40248119

Summary

A hollow triangular rod-type valley-Hall photonic topological insulator is proposed, and two tandem residual deep neural networks are built for multimodal inverse design of the structure. One of them is a tandem multilayer perceptron, and the other is a composite tandem network based on variational auto-encoder. The former is used to inversely infer the value of the structural sizes, and the latter is used to predict the structural image of the lattice from demanded design targets. Residual connections are included in both networks to speed up the training convergence as well as avoid vanishing gradient problem. Based on an arbitrary inversely designed lattice, domain walls between two photonic topological insulators with different topology are constructed, and full-wave simulations on the transmission properties are conducted. Numerical results show that robust topologically protected wave propagation is supported along the domain wall with little backscattering, demonstrating that the proposed methods are valid.

Subject areas: Topological photonics, Deep learning, Inverse design

Graphical abstract

graphic file with name fx1.jpg

Highlights

  • •

    A valley-Hall photonic topological insulator composed of hollow triangular rods is proposed

  • •

    A tandem residual multilayer perceptron is built for inverse design of the structural size

  • •

    A composite tandem encoder network is established for inverse design of the structural image


Topological photonics; Deep learning; Inverse design

Introduction

With the discovery of the photonic quantum Hall effects,1,2,3,4,5 it has been found that research on photonic topological insulators (PTIs) is of great significance. PTI6,7,8,9,10 in the photonic system can mimic the quantum Hall effect in the electronic system with the absence of the magnetic field, thereby allowing topologically protected edge states. These edge states have been verified by experiments and applied to various devices such as topological lasers,11 optical waveguides,12,13,14 modulators,15,16 and communication components.17,18 The key factor affecting the electromagnetic (EM) characteristics of these devices with artificial periodic structures is the structural parameters. In practical applications, designers often need to conduct inverse designs starting from the final goals of the EM responses to obtain suitable structural sizes. Currently, several methods have been reported to solve this inverse design problems in photonic fields, including many popular algorithms such as topology optimization,19,20,21 automatic differentiation,22,23 genetic algorithm,24,25 simulated annealing,26 and particle swarm optimization.27,28 These methods touch optimal results basically by continuous updates and require repetitive iterations for different optimization objectives. They sometimes may not have enough high efficiency and are not good at handling complex high-dimensional data relationships. In recent years, deep learning methods have provided a new data-driven approach for the inverse EM problems.29,30,31,32,33,34,35,36 In topological photonic field, well-trained deep neural networks have attained good performance in realizing various functionality.37,38,39 This method can significantly reduce the time consumption in the searching process, which is of great importance to the design of photonic band structures and topological phases. Nevertheless, there comes out new issues that for many situations different photonic structures may have almost the same EM response, giving rise to prediction error by conventional neural networks. To deal with the non-uniqueness issue, a tandem neural network is proposed that trains the inverse network along with a pre-trained forward network.40 One may also face the vanishing gradient problem during the training. Thus, researchers improved the tandem network by different methods including modifying the loss functions41 or using an adaptive updating strategy.42 Later, a tandem residual convolutional network for the inverse design of multiplexed supercell metasurface is reported.43 Another double-check tandem network with two forward models44 is presented to avoid gradient vanishing and keep the predicted structure in the valid range.

In this study, we propose a valley-hall PTI formed by hexagonal-lattice hollow triangular rods. A tandem multilayer perceptron (MLP) and a composite tandem variational auto-encoder (VAE) network with residual blocks are applied to the inverse design of the PTI. Tandem networks are used to solve “one to many” problem and avoid invalid designs in both networks. Residual connections are utilized in the fully connected layers and convolutional layers to prevent the vanishing gradient problem existing in traditional tandem networks, which improves the convergence speed and the accuracy of the prediction. We further demonstrate the validity of the predicted structure by finite element method (FEM) simulations on the band structures and edge waveguiding for an arbitrary input demand. This work provides improved deep-learning-based methods for multimodal inverse design of PTIs, which can also be extended to other photonic devices such as metasurfaces.

Results

Models of the valley-Hall PTI

The proposed two-dimensional PTI consists of hollow triangular rods in the center of each unit cell arranged in hexagonal lattices, as shown in Figure 1A. The lattice constant is a, and the dielectric constant of the hollow triangular rod in the air background is ε. The side length of the outer and inner triangles divided by a is d1 and d2, respectively. The triangular rod has a rotation angle θ starting from the positive y axis direction, which can be adjusted for required configurations in practical use. In general, in this unit cell, the four parameters (d1, d2, θ, and ε) form the design space. By scanning along the edge of the first Brillouin zone with 29 sampled points through FEM calculation, we in Figure 1B present the first six transverse electric (TE) photonic bands (in-plane Hz≠0) of a typical case (d1 = 0.60, d2 = 0.16, θ = 0°, and ε = 14) of the PTI. Obviously, there is a photonic band gap (PBG) marked with PBG 1 between Band 2 and Band 3. The center frequency f and gap ratio g of PBG 1 are defined by:

f=wx+wy2,g=2(wx−wy)wx+wy×100%

where wx is the minimum of Band 3, and wy is the maximum of Band 2. f is expressed in units of c/a. It is found that when we increase the rotation angle θ and keep other parameters fixed, PBG 1 will close at θ = 30° and form a Dirac point as marked in red in Figure 1C. After θ rises over 30°, PBG 1 reopens, and the bands are almost the same as those with the rotation angle of 60°−θ. Until θ = 60°, PBG 1 returns to the same state at θ = 0° in Figure 1B because of the symmetry. This PBG 1 shows topological valley-Hall phase, which can be characterized by the valley Chern numbers45:

CK∕K′=(2π)−1∫HBZFkⅆSk

where the integral is over half of the first Brillouin zone around the K or K′ points and the Berry curvature is expressed as Fk=∇k×Ak=∇k×⟨uk|∇k|uk⟩.46 With the reported computation method,47 we calculated the valley Chern number and obtained non-vanishing CK>0 (or CK′<0). As a result, PBG 1 exhibits topologically non-trivial phase, and topologically protected edge states can be expected within the band gap. This is the reason we choose PBG 1 as the target band gap in the following inverse design tasks. Here, we prefer a hollow triangular rod to a solid triangular rod in the proposed PTI mainly because the array of solid triangular rod with the same side length has narrower PBG, which is presented in Figure S1 and Table S1 of the supplemental information.

Figure 1.

Figure 1

The structure of the proposed PTI and its photonic bands

(A) The lattice structure of the hollow triangular rod-type PTI.

(B) The band structure at θ = 0° or θ = 60° with PBG 1 appears.

(C) The band structure at θ = 30° with a Dirac point marked in red.

Inverse design based on tandem MLP network (Method 1)

For inverse design of the hollow triangular rod-type PTI through deep-learning-based methods, we build two deep neural networks in different purpose. The first is a tandem MLP named as Method 1. This method is intended to predict a practicable value set (d1, d2, θ, ε) from an on-demand target of the center frequency and gap ratio (f, g) of the PBG, rather than from the photonic bands. The reason lies in that it is much easier for a designer to raise a required (f, g) than the photonic bands. Considering the truth that lattices with symmetric structure may have the same bands, through fast loop of FEM calculation, we rapidly generate 6,000 data samples of the bands and (f, g) for 6,000 groups of (d1, d2, θ, ε) as the dataset, where d1∈[0.10,0.80], d2∈[0.08,0.78], θ∈[0°, 60°], and ε∈[14, 30]. The tandem MLP consists of two models: Model A and Model B, as shown in Figure 2. Model A is used to fit the mapping function between the two properties (f, g) of PBG 1 and the two bands forming PBG 1 that totally contains 58 (29×2) eigenfrequencies (ω1, ω2, …,ω58). This model is implemented with 5 fully connected layers having the neurons 2-64-128-64-58. Except the last layer, each layer is followed by the widely used activation function ReLU. Adam is chosen as the optimizer, and its learning rate is set to 3E−4. The loss is evaluated by mean squared error (MSE), and the batch size is set to 32. The prepared data are randomly divided into training set and test set in a ratio of 4:1. For the evaluation of the network’s performance, the accuracy is defined as

Accuracy=1-∑i=1N|yi−yˆi|N·[max(y)−min⁡(y)]

where N is the size of the dataset, yi is the ground truth, and yˆi is the predicted value. Since in this work y value are normalized to [−1,1], max(y) = 1 and min⁡(y) = −1. With the training for 500 epochs, the prediction accuracy on the test set reaches 98.10%.

Figure 2.

Figure 2

The architecture of the proposed network in Method 1

Model A consists of single MLP from the center frequency and gap ratio (f,g) to the 58 eigenfrequencies of two target bands. Model B consists of tandem resisual MLPs from the eigenfrequencies to the structural parameters (d1, d2, θ, ε) then back to the eigenfrequencies.

After the two bands related to PBG 1 are inferred from Model A, the next-step Model B takes the 58 eigenfrequencies (ω1,ω2, …,ω58) of the two bands as the inputs to predict the parameters (d1, d2, θ, and ε) of the PTI. Since two cells with different structural parameters may have almost the same band structure, Model B faces the “one to many” problem that there are multiple outputs corresponding to the same input. Under such situation, an ordinary MLP network may finally output an average of these outputs, leading to low accuracy of the results. For this reason, we use tandem network composed of a forward model and an inverse model: the forward model outputs the bands by inputting the structural parameters, and the inverse model has just the opposite function with its outputs connected to the inputs of the forward model. As the tandem model learns by comparing the ground truth bands to the predicted bands, the inverse model will only need to find one of the many possible structures that can have the correct bands. Here, the forward model is implemented by a 7-layer MLP with the neurons 4-16-32-128-256-128-58, while the inverse model is implemented by another residual MLP (ResMLP) consisting of one 58-neuron input layer, three residual blocks, and one 4-neuron output layer. Since all data will be linearly normalized to (−1, 1), tanh function is used after the last layer of the inverse model as the activation function. It limits the range of the outputs to make 100% sure that the predicted structure is always valid and physically realizable. Residual blocks are added to overcome the vanishing gradient issue, and each residual block contains three 128-neuron fully connected layers. All other activation function after each layer is ReLU. The optimizer is Adam with the learning rate of 3E−4. The loss still uses MSE and the batch size is 32. To train the tandem model, 80% of the prepared data are taken for training and the rest 20% are taken for testing. The forward model is trained firstly for 500 epochs, and its accuracy on test set reaches 99.65%. Then the inverse model updates its parameters in training based on the well-trained forward model with fixed parameters. In order to better highlight the advantages of the proposed tandem ResMLP, we also trained traditional tandem network40 and double-check network44 of the same network depth, respectively, for the same task. As compared in Figure 3, for either the training set or the test set, the MSE loss of the tandem ResMLP drops much faster and finally reaches lower than that of the other two during the process of 500 epochs. The final accuracy of the inverse model and the total network with the three different model types for the test set are compared in Table 1. These two accuracies of the tandem ResMLP reaches 99.70% and 98.85%, respectively, which are the highest of the three networks. Therefore, the tandem ResMLP is able to outperform the other two tandem networks in the inverse design task of the PTI structural size.

Figure 3.

Figure 3

The MSE loss of three different models under 500 epochs of the training

The loss curves of traditional tandem model, double-check tandem model, and the tandem ResMLP are drawn in black, red, and blue, respectively.

(A) The loss curves of the three models on training set.

(B) The loss curves of the three models on test set.

Table 1.

Performance of three different model types

Model type Accuracy of the inverse model Accuracy of the total network
Traditional tandem 99.35% 97.72%
Double-check tandem 99.56% 98.22%
Tandem ResMLP 99.70% 98.85%

Inverse design based on composite VAE network (Method 2)

Secondly, a composite network constituted by VAE and ResMLP is proposed as Method 2 for multimodal inverse design of the PTI. It purposes to reconstruct the sample image of the PTI lattice by the given center frequency and gap ratio (f, g), whose workflow is divided into three models presented in Figure 4A. Model A is an MLP predicting the two bands related to PBG 1 from (f, g), which is completely the same as the Model A in Method 1. Model B is another MLP taking the eigenfrequencies of the two bands to infer the latent features of the sample image. Model C is the decoder part of a VAE, which reconstructs the image of the lattice from the latent features. Since there is no need to train Model A again, for Model B and C, the key step is to obtain a well-trained VAE, whose construction is illustrated in Figure 5. The encoder takes a one-channel grayscale image of the PTI lattice with 128 by 128 pixel as the input. Here, the structural parameters d1, d2, and θ are directly displayed on the image, while the dielectric constant ε is mapped to the grayscale of the image. The image is then processed through 6 convolutional layers with two residual connections as well as linear layers to form the latent space with a normal distribution. Four-dimensional latent feature vector (z1, z2, z3, z4) is extracted by sampling in the normal distribution. At last, the decoder recovers the 128 by 128 pixel 1-channel image of the lattice from the latent vector through 6 transposed convolutional layers with two residual connections. Based on the prepared 6,000 data samples mentioned in Method 1, we generated about 18,000 images for this training by data argumentation including horizontal and vertical flip according to the symmetry of the lattice. They are divided into training set and test set in a ratio of 4:1. The setting of the optimizer, learning rate, and batch size are the same as in Method 1. The loss is calculated using the sum of binary cross entropy and Kullback–Leibler divergence. The mean absolute error (MAE) is used to evaluate the accuracy on the test set, which is defined as the average grayscale difference between the reconstructed image and the original image pixel by pixel. After 100-epoch training, the MAE on the test set drops to 3.2E−4, proving that this VAE has successfully built a mapping between the image space and the latent space.

Figure 4.

Figure 4

The architecture of the proposed composite network in Method 2

(A) The inverse design workflow from the center frequency and gap ratio (f, g) to the sample images.

(B) The verification workflow from the designed images to the center frequency and gap ratio (f, g).

Figure 5.

Figure 5

The proposed VAE network for realizing Model C of the composite network

(A) The architecture of the VAE including the Encoder Q and the Decoder P. The input layer is represented as 128×128×1 that stands for 128 by 128 pixel and 1 channel. All other layers have the similar representaition.

(B) The loss curves of the VAE on the training set and test set under 100 epochs of the training.

The last step is to obtain Model B that can deduce the latent feature vector (z1, z2, z3, z4) of an image from the given bands (ω1,ω2, …,ω58). As mentioned before, the images with symmetric structure have the same bands, which means one given band set may correspond to several latent vectors. Therefore, this is also a “one to many” problem where we should not simply assess the loss by comparing the ground truth image to the predicted image. Based on the previous experience, we adopt the forward model MLP 1 (latent vector to bands) and the inverse model MLP 2 (bands to latent vector) illustrated in Figure 6 to compose a similar tandem ResMLP network with that in Method 1. With the prepared 6,000 pairs of the sample image and its bands, the Encoder Q of the VAE helps to generate sampled latent vector of each pair. Under the same hyper-parameters as in Method 1, MLP 1 is trained firstly for 400 epochs, and its accuracy on test set is 99.31%. MLP 2 is trained secondly based on the well-trained MLP 1 with fixed parameters, and its accuracy on test set is 99.64% after 400 epochs. Finally, we concatenate Model A, MLP 2 as Model B, and Decoder P as Model C to finish the composite network in Figure 4A for the inverse design from (f, g) to the image of the lattice. The predicted image is processed through Encoder Q cascaded with MLP 1 to get corresponding bands and (f, g) for verification, as shown in Figure 4B. The accuracy of the predicted bands on the test set reaches 98.6%, and the accuracy for (f, g) reaches 97.5%. The image can also be imported into FEM software (Comsol Multiphysics) to calculate the bands for comparison with the target bands (output of Model A) and required (f, g). It is worth mentioning that the large dataset of 6,000 samples is not strictly necessary. We also trained the neural networks in the two methods using a smaller dataset (containing 1,500 samples) for the inverse design tasks. Results show that the smaller dataset achieves the accuracy of 98.2% and 97.5% of the predicted bands by Method 1 and Method 2, respectively, as well as the accuracy of 97.7% and 96.9% of the center frequency and gap ratio. These results are compared with those trained by the large dataset in Table S2 of the supplemental information, which demonstrates that the proposed methods can still obtain acceptable performance with fewer training data.

Figure 6.

Figure 6

The proposed tandem MLP network for realizing Model B of the composite network

(A) The architecture of the tandem network including the forward (MLP 1) and the inverse (MLP 2) model.

(B) The loss curves of MLP 2 on the training set and test set under 400 epochs of the training.

Verification on inversely designed topologically protected transport

To demonstrate the effectiveness of the proposed inverse design methods, we use Method 1 and Method 2, respectively, to carry out a same task of designing a valley-Hall PTI with on-demand PBG. We arbitrarily raise a requirement of the target center frequency 0.48 THz and gap ratio 25%, which is (f, g) = (0.48, 25%) when we take the lattice constant a = 0.3 mm. Through the first step, model A predicts the target bands as shown in blue curve of Figure 7A. Method 1 obtains the inversely designed structural parameters (d1 = 0.80, d2 = 0.45, θ = 8.65°, and ε = 23.87), while Method 2 obtains the image of the lattice. The designed lattice structures by the two methods are compared in the inset of Figure 7A. The calculated bands of the two structures by FEM, respectively, drawn in black and red curve illustrate that they are very close to the blue target bands. The PBG by Method 1 has a center frequency and gap ratio of (0.478, 25.1%), while that by Method 2 has (0.479, 24.9%). Both results are substantially consistent with the raised design requirement. Although this study is based on a fixed triangular geometry, the proposed methods are generalizable. To demonstrate this, we applied our methods to perform multimodal inverse design on a ring-shaped spin-Hall PTI. The new dataset contains 2,000 samples, and we toke several tens of hours to fabricate it. Using Method 1 and Method 2 for the inverse design of this spin-Hall PTI, the accuracy of the predicted band structure reaches 98.1% and 97.6%, respectively. For example, we selected a representative set of the target central frequency and band-gap ratio (0.400, 20%) and input them into the two neural networks. The lattice structures and band structures obtained through inverse design using Method 1 and Method 2 are shown in Figure S2 of the supplemental information. As observed, the band structures derived from both methods closely match the target band structure. Specifically, the central frequency and band-gap ratio achieved by Method 1 are (0.400, 19.4%), while those achieved by Method 2 are (0.413, 20.5%), both meeting the design objectives. These results prove that the proposed methods can be efficiently extended to the design of other types of PTIs.

Figure 7.

Figure 7

Inverse design verification of a raised requirement (f, g) = (0.48, 25%)

(A) Comparison of the target bands with the derived bands from two predicted structures. The markers on each curve represent the 29 sampled frequency points of each band.

(B) The band structure of the designed PTI and the field profiles of the LCP and RCP modes at the K and K′ valleys.

Furthermore, the inversely designed PTI structure is used for constructing topologically protected edge waveguides. For simplicity, we take the inversely designed result of Method 1 as an example. The derived topologically non-trivial PBG ranges from 0.418 c/a to 0.538 c/a, as highlighted in Figure 7B. The spiral field profiles of the left circularly polarized (LCP) and right circularly polarized (RCP) modes at the K(K′) valleys are observed in the inset, revealing that the inter-valley coupling between the K and K′ valleys will be suppressed. It should be noted that, in this case, another PBG emerges between the first and second bands. However, this PBG presents a topologically trivial phase, thus is not in our consideration of the design target. To verify the existence of the topologically protected edge states within the target non-trivial PBG, we constructed a supercell composed of two different PTIs opposite with valley Chern numbers in Figure 8A: the upper PTI has the parameters (d1 = 0.80, d2 = 0.45, θ = 8.65°, and ε = 23.87), while the lower one has (d1 = 0.80, d2 = 0.45, θ′ = 51.35°, and ε = 23.87) satisfying θ′ = 60°−θ. By applying periodic boundary condition on the x direction, we calculated the projected band diagram of the supercell as shown in Figure 8B. It is obvious that one edge state appears as drawn in red curve within PBG 1, which is protected by the different topology between the two PTIs. This state covers the whole frequency range of PBG 1 so the width of the high transmission working band is expected to be equal to that of PBG 1. The electric field distribution in Figure 8C exhibits good confinement of the edge state at the interface. In addition, another band gap is observed below PBG 1 with no edge states, which agrees with previous analysis and further verifies the topologically trivial phase of this band gap.

Figure 8.

Figure 8

Topologically protected edge states at a zigzag interface between two PTIs with different topology

(A) A supercell composed of upper PTI with θ = 8.65° and lower PTI with θ′ = 51.35°.

(B) The projected bands of the supercell. The red line is the edge state in PBG 1.

(C) The field distribution of the edge state at kx = 0.

In the end, the topological protected edge wave propagation along the domain walls formed by inversely designed cells is studied. We constructed a straight domain wall and a non-straight domain wall with a 120° turn by neighboring the two different PTIs. With the excitation from a point source at the entrance, the transmission spectra through the straight and non-straight edge channels are calculated by full-wave simulation. As shown in Figure 9, both channels have very high transmittance larger than 98% within the working band from 0.42 c/a to 0.53 c/a, which is 0.42 to 0.53 THz under a = 0.3 mm. No matter the propagation has a 120° turn or not, the field distributions of the propagating modes manifest highly steady status, revealing that the backscattering is greatly suppressed because of the photonic valley-Hall effect. These results are in agreement with the initial design requirement of the center frequency and gap ratio.

Figure 9.

Figure 9

Topological protected edge wave propagation in the domain walls with inversely deigned lattice

(A) The field profiles of the topological transport through a straight domain wall.

(B) The field profiles of the topological transport through a non-straight domain wall with a 120° turn.

(C) The transmission spectrum of the straight domain wall in the frequency range of 0.35∼0.60 c/a.

(D) The transmission spectrum of the domain wall with a 120° turn in the frequency range of 0.35∼0.60 c/a.

Discussion

In summary, a hollow triangular rod-type PTI showing valley-Hall phase is proposed, and two tandem neural networks with residual blocks are presented for multimodal inverse design of the PTI. In the inverse models of both the tandem MLP and composite VAE networks, residual connections work in with the tanh activation function. The former helps to accelerate the convergence and prevent the gradient vanishing, while the latter gives the assurance of absolute validity of the inversely designed structures. Compared with reported traditional tandem network and double-check network, the proposed tandem residual networks manifest better performance on the training speed and prediction accuracy as well as provide multimodal results including the parameters and the images in the “one to many” design tasks. By using the well-trained network, we completed a trial inverse design on the topological edge waveguides according to an arbitrary design requirement. Full-wave simulation results indicate that the robust topologically protected edge propagation is supported with little backscattering even through sharp turns, illustrating that the proposed inverse design methods based on deep learning are effective and can be used in realizing low-loss waveguides and more photonic devices.

Limitations of the study

We intend to use deep-learning-based method instead of traditional numerical computation to efficiently design a PTI meeting a specific demand on the working band. The optimization of structural parameters for objectives like maximizing the band gap is not studied because in certain cases narrowband PTI devices are demanded. When extended to other photonic structures, the proposed methods require new datasets for training. Although the data can be produced fast with the calculation technology nowadays, the generality of this method is limited to some degree. Moreover, the inversely designed topological waveguides need further experimental demonstration to prove its practical effectiveness.

Resource availability

Lead contact

Further information and requests for resources should be directed to and will be fulfilled by the lead contact, Le Zhang (zhangle@cjlu.edu.cn).

Materials availability

This study did not generate new unique reagents.

Data and code availability

  • •

    Data reported in this paper will be shared by the lead contact upon request.

  • •

    This paper does not report original code.

  • •

    Any additional information required to reanalyze the data reported in this paper is available from the lead contact upon request.

Acknowledgments

This work was financially supported by the National Natural Science Foundation of China (62175223) and Key Research and Development Program of Zhejiang Province (2024C01028, 2024C01102, 2024C01108, and 2024C01106).

Author contributions

Conceptualization, L.Z.; methodology, L.Z. and B.-X.W.; investigation, B.-J.W., B.-X.W., and Y.-G.X.; writing – original draft, B.-J.W. and L.Z.; writing – review and editing, L.Z. and B.-X.W.; funding acquisition, L.Z. and D.-P.Z.; supervision, J.-H.C.

Declaration of interests

The authors declare no competing interests.

STAR★Methods

Key resources table

REAGENT or RESOURCE SOURCE IDENTIFIER
Software and algorithms

Comsol Multiphysics COMSOL Co., Ltd. https://www.comsol.com/products
PyTorch The PyTorch Foundation https://pytorch.org/

Method details

The finite element method through the software Comsol Multiphysics 5.6 was used in all the simulations. The Eigenfrequency Solver was applied in calculating the photonic band structure and periodic conditions are set at each pair of the opposite boundaries of the unit cell. Frequency Domain Solver was applied for simulations on the edge propagating behavior of the proposed photonic topological insulator. Point wave excitation with a frequency of 0.47 c/a was set at the input port of the domain wall channels.

The numerical range of the structural parameters is predefined as d1∈[0.10,0.80], d2∈[0.08,0.78], θ∈[0°, 60°], ε∈[14, 30]. For the training of proposed models, within this range 6,000 data samples of the bands with center frequency and gap ratio (f, g) were generated for 6000 groups of the structural parameters (d1, d2, θ, ε) and corresponding lattice image as the dataset. All the deep-learning models are constructed on Pytorch as described in the results section of the main text. For the tandem MLP networks in both Method 1 and Method 2, we set the learning rate as 3E−4, batch size as 32, MSE as the loss function and train the network for 400∼500 epochs. For the VAE network in Method 2, we set the learning rate as 5E−4, batch size as 32, BCE+KLD as the loss function and train the network for 100 epochs.

Quantification and statistical analysis

The simulation data were calculated by Comsol Multiphysics software. Figures in the main text were produced by Post-processing of Comsol and Python Matplotlib package.

Published: March 22, 2025

Footnotes

Supplemental information can be found online at https://doi.org/10.1016/j.isci.2025.112276.

Contributor Information

Le Zhang, Email: zhangle@cjlu.edu.cn.

Ben-Xin Wang, Email: wangbenxin@jiangnan.edu.cn.

Supplemental information

Document S1. Figures S1 and S2 and Tables S1 and S2
mmc1.pdf (297.9KB, pdf)

References

  • 1.Kane C.L., Mele E.J. Quantum spin Hall effect in graphene. Phys. Rev. Lett. 2005;95 doi: 10.3322/caac.21660. [DOI] [PubMed] [Google Scholar]
  • 2.Young A.F., Sanchez-Yamagishi J.D., Hunt B., Choi S.H., Watanabe K., Taniguchi T., Ashoori R.C., Jarillo-Herrero P. Tunable symmetry breaking and helical edge transport in a graphene quantum spin Hall state. Nature. 2014;505:528–532. doi: 10.1038/nature12800. [DOI] [PubMed] [Google Scholar]
  • 3.Ma T., Shvets G. All-Si valley-Hall photonic topological insulator. New J. Phys. 2016;18 doi: 10.1088/1367-2630/18/2/025012. [DOI] [Google Scholar]
  • 4.Wang Y. Anomalous quantum hall effect of light in bloch-wave modulated photonic crystals. Phys. Rev. Lett. 2019;122 doi: 10.1103/PhysRevLett.122.233904. [DOI] [PubMed] [Google Scholar]
  • 5.Xie B., Su G., Wang H.F., Liu F., Hu L., Yu S.Y., Zhan P., Lu M.H., Wang Z., Chen Y.F. Higher-order quantum spin Hall effect in a photonic crystal. Nat. Commun. 2020;11:3768. doi: 10.1038/s41467-020-17593-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Barik S., Miyake H., DeGottardi W., Waks E., Hafezi M. Two-dimensionally confined topological edge states in photonic crystals. New J. Phys. 2016;18 doi: 10.1088/1367-2630/18/11/113013. [DOI] [Google Scholar]
  • 7.Wu L.H., Hu X. Scheme for achieving a topological photonic crystal by using dielectric material. Phys. Rev. Lett. 2015;114 doi: 10.1103/PhysRevLett.114.223901. [DOI] [PubMed] [Google Scholar]
  • 8.Anderson P.D., Subramania G. Unidirectional edge states in topological honeycomb-lattice membrane photonic crystals. Opt. Express. 2017;25:23293–23301. doi: 10.1364/OE.25.023293. [DOI] [PubMed] [Google Scholar]
  • 9.Wang Q., Xue H., Zhang B., Chong Y.D. Observation of protected photonic edge states induced by real-space topological lattice defects. Phys. Rev. Lett. 2020;124 doi: 10.1103/PhysRevLett.124.243602. [DOI] [PubMed] [Google Scholar]
  • 10.Yang Y., Xu Y.F., Xu T., Wang H.X., Jiang J.H., Hu X., Hang Z.H. Visualization of a unidirectional electromagnetic waveguide using topological photonic crystals made of dielectric materials. Phys. Rev. Lett. 2018;120 doi: 10.1103/PhysRevLett.120.217401. [DOI] [PubMed] [Google Scholar]
  • 11.Zeng Y., Chattopadhyay U., Zhu B., Qiang B., Li J., Jin Y., Li L., Davies A.G., Linfield E.H., Zhang B., et al. Electrically pumped topological laser with valley edge modes. Nature. 2020;578:246–250. doi: 10.1038/s41586-020-1981-x. [DOI] [PubMed] [Google Scholar]
  • 12.He X.T., Liang E.T., Yuan J.J., Qiu H.Y., Chen X.D., Zhao F.L., Dong J.W. A silicon-on-insulator slab for topological valley transport. Nat. Commun. 2019;10:872. doi: 10.1038/s41467-019-08881-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.Noh J., Benalcazar W.A., Huang S., Collins M.J., Chen K.P., Hughes T.L., Rechtsman M.C. Topological protection of photonic mid-gap defect modes. Nat. Photonics. 2018;12:408–415. doi: 10.1038/s41566-018-0179-3. [DOI] [Google Scholar]
  • 14.Kuruma K., Yoshimi H., Ota Y., Katsumi R., Kakuda M., Arakawa Y., Iwamoto S. Topologically-protected single-photon sources with topological slow light photonic crystal waveguides. Laser Photon. Rev. 2022;16 doi: 10.1002/lpor.202200077. [DOI] [Google Scholar]
  • 15.Shalaev M.I., Walasik W., Litchinitser N.M. Optically tunable topological photonic crystal. Optica. 2019;6:839–844. doi: 10.1364/OPTICA.6.000839. [DOI] [Google Scholar]
  • 16.Tang G.J., He X.T., Shi F.L., Liu J.W., Chen X.D., Dong J.W. Topological photonic crystals: physics, designs, and applications. Laser Photon. Rev. 2022;16 doi: 10.1002/lpor.202100300. [DOI] [Google Scholar]
  • 17.Shalaev M.I., Walasik W., Tsukernik A., Xu Y., Litchinitser N.M. Robust topologically protected transport in photonic crystals at telecommunication wavelengths. Nat. Nanotechnol. 2019;14:31–34. doi: 10.1038/s41565-018-0297-6. [DOI] [PubMed] [Google Scholar]
  • 18.Yang Y., Yamagami Y., Yu X., Pitchappa P., Webber J., Zhang B., Fujita M., Nagatsuma T., Singh R. Terahertz topological photonics for on-chip communication. Nat. Photonics. 2020;14:446–451. doi: 10.1038/s41566-020-0618-9. [DOI] [Google Scholar]
  • 19.Christiansen R.E., Wang F., Sigmund O., Stobbe S. Designing photonic topological insulators with quantum-spin-Hall edge states using topology optimization. Nanophotonics. 2019;8:1363–1369. doi: 10.1515/nanoph-2019-0057. [DOI] [Google Scholar]
  • 20.Chen Y., Meng F., Li G., Huang X. Topology optimization of photonic crystals with exotic properties resulting from Dirac-like cones. Acta Mater. 2019;164:377–389. doi: 10.1016/j.actamat.2018.10.058. [DOI] [Google Scholar]
  • 21.Meng F., Jia B., Huang X. Topology-Optimized 3D Photonic Structures with Maximal Omnidirectional Bandgaps. Adv. Theory Simul. 2018;1 doi: 10.1002/adts.201800122. [DOI] [Google Scholar]
  • 22.Wang H., Guo C., Zhao Z., Fan S. Compact incoherent image differentiation with nanophotonic structures. ACS Photonics. 2020;7:338–343. doi: 10.1029/2021WR030595. [DOI] [Google Scholar]
  • 23.Colburn S., Majumdar A. Inverse design and flexible parameterization of meta-optics using algorithmic differentiation. Commun. Phys. 2021;4:65. doi: 10.1038/s42005-021-00568-6. [DOI] [Google Scholar]
  • 24.Ren Y., Zhang L., Wang W., Wang X., Lei Y., Xue Y., Sun X., Zhang W. Genetic-algorithm-based deep neural networks for highly efficient photonic device design. Photonics Res. 2021;9:B247–B252. doi: 10.1364/PRJ.416294. [DOI] [Google Scholar]
  • 25.Ma Y., Ferguson A.L. Inverse design of self-assembling colloidal crystals with omni-directional photonic bandgaps. Soft Matter. 2019;15:8808–8826. doi: 10.1039/C9SM01500K. [DOI] [PubMed] [Google Scholar]
  • 26.Liao Y., Yu T., Wang Y., Dong B., Yang G. Simulated annealing algorithm with neural network for designing topological photonic crystals. Opt. Express. 2023;31:31597–31609. doi: 10.1364/OE.500720. [DOI] [PubMed] [Google Scholar]
  • 27.Mirjalili S.M., Mirjalili S., Lewis A., Abedi K. A tri-objective particle swarm optimizer for designing line defect photonic crystal waveguides. Photonics Nanostruct. 2014;12:152–163. doi: 10.1016/j.photonics.2013.11.001. [DOI] [Google Scholar]
  • 28.Wang Y.H., Zhang H.F. Angular insensitive nonreciprocal ultrawide band absorption in plasma-embedded photonic crystals designed with improved particle swarm optimization algorithm. Chinese Phys. B. 2023;32 doi: 10.1088/1674-1056/ac8929. [DOI] [Google Scholar]
  • 29.Ma W., Cheng F., Liu Y. Deep-learning-enabled on-demand design of chiral metamaterials. ACS Nano. 2018;12:6326–6334. doi: 10.1021/acsnano.8b03569. [DOI] [PubMed] [Google Scholar]
  • 30.Liu Z., Zhu D., Rodrigues S.P., Lee K.T., Cai W. Generative model for the inverse design of metasurfaces. Nano Lett. 2018;18:6570–6576. doi: 10.1021/acs.nanolett.8b03171. [DOI] [PubMed] [Google Scholar]
  • 31.da Silva Ferreira A., Malheiros-Silveira G.N., Hernández-Figueroa H.E. Comput-ing optical properties of photonic crystals by using multilayer perceptron and extreme learning machine. J. Lightwave Technol. 2018;36:4066–4073. doi: 10.1109/JLT.2018.2856364. [DOI] [Google Scholar]
  • 32.Asano T., Noda S. Optimization of photonic crystal nanocavities based on deep learning. Opt. Express. 2018;26:32704–32717. doi: 10.1364/OE.26.032704. [DOI] [PubMed] [Google Scholar]
  • 33.Li X., Ning S., Liu Z., Yan Z., Luo C., Zhuang Z. Designing photonic crystal with anticipated band gap through a deep learning based data-driven method. Comput. Methods Appl. Mech. Eng. 2020;361 doi: 10.1016/j.cma.2019.112737. [DOI] [Google Scholar]
  • 34.Liu Z., Raju L., Zhu D., Cai W. A hybrid strategy for the discovery and design of photonic structures. IEEE. J. Em. Sel. Top. C. 2020;10:126–135. doi: 10.1109/JETCAS.2020.2970080. [DOI] [Google Scholar]
  • 35.Jiang J., Fan J.A. Multiobjective and categorical global optimization of photonic structures based on ResNet generative neural networks. Nanophotonics. 2020;10:361–369. doi: 10.1515/nanoph-2020-0407. [DOI] [Google Scholar]
  • 36.Liu Z., Zhu D., Raju L., Cai W. Tackling photonic inverse design with machine learning. Adv. Sci. 2021;8 doi: 10.1002/advs.202002923. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 37.Christensen T., Loh C., Picek S., Jakobović D., Jing L., Fisher S., Ceperic V., Joannopoulos J.D., Soljačić M. Predictive and generative machine learning models for photonic crystals. Nanophotonics. 2020;9:4183–4192. doi: 10.1515/nanoph-2020-0197. [DOI] [Google Scholar]
  • 38.Singh R., Agarwal A., W Anthony B. Mapping the design space of photonic topological states via deep learning. Opt. Express. 2020;28:27893–27902. doi: 10.1364/OE.398926. [DOI] [PubMed] [Google Scholar]
  • 39.He L., Wen Z., Jin Y., Torrent D., Zhuang X., Rabczuk T. Inverse design of topological metaplates for flexural waves with machine learning. Mater. Des. 2021;199 doi: 10.1016/j.matdes.2020.109390. [DOI] [Google Scholar]
  • 40.Liu D., Tan Y., Khoram E., Yu Z. Training deep neural networks for the inverse desi-gn of nanophotonic structures. ACS Photonics. 2018;5:1365–1369. doi: 10.1021/acsphotonics.7b01377. [DOI] [Google Scholar]
  • 41.Xu X., Sun C., Li Y., Zhao J., Han J., Huang W. An improved tandem neural netw-ork for the inverse design of nanophotonics devices. Opt. Commun. 2021;481 doi: 10.1016/j.optcom.2020.126513. [DOI] [Google Scholar]
  • 42.Chen J., Dai Z., Yang Z., Pan Y., Zhang X., Wu J., Reza Soltanian M. An improved tandem neural network architecture for inverse modeling of multicomponent reactive transport in porous media. Water Resour. Res. 2021;57 doi: 10.1029/2021WR030595. [DOI] [Google Scholar]
  • 43.Yeung C., Tsai J.M., King B., Pham B., Ho D., Liang J., Knight M.W., Raman A.P. Multiplexed supercell metasurface design and optimization with tandem residual networks. Nanophotonics. 2021;10:1133–1143. doi: 10.1515/nanoph-2020-0549. [DOI] [Google Scholar]
  • 44.Gu Y., Li D., Li E.P. 2022 International Conference on Microwave and Millimeter Wave Technology (ICMMT) IEEE; 2022. Double-Check Tandem Neural Network for Inverse Design of Photonic Band Structures; pp. 1–3. [DOI] [Google Scholar]
  • 45.Yang Y., Xu Z., Sheng L., Wang B., Xing D.Y., Sheng D.N. Time reversal symmetry broken quantum spin Hall effect. Phys. Rev. Lett. 2011;107 doi: 10.1103/PhysRevLett.107.066602. [DOI] [PubMed] [Google Scholar]
  • 46.Xiao D., Chang M.C., Niu Q. Berry phase effects on electronic properties. Rev. Mod. Phys. 2010;82:1959–2007. doi: 10.1103/RevModPhys.82.1959. [DOI] [Google Scholar]
  • 47.Fukui T., Hatsugai Y., Suzuki H. Chern numbers in discretized Brillouin zone: effici-ent method of computing (spin) Hall conductances. J. Phys. Soc. Jpn. 2005;74:1674–1677. doi: 10.1143/JPSJ.74.1674. [DOI] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Document S1. Figures S1 and S2 and Tables S1 and S2
mmc1.pdf (297.9KB, pdf)

Data Availability Statement

  • •

    Data reported in this paper will be shared by the lead contact upon request.

  • •

    This paper does not report original code.

  • •

    Any additional information required to reanalyze the data reported in this paper is available from the lead contact upon request.


Articles from iScience are provided here courtesy of Elsevier

RESOURCES