Abstract
We review simulation optimization methods and their connection with artificial intelligence (AI) techniques. In particular, we focus on two areas: stochastic gradient estimation, which plays a central role in training neural networks for deep learning and reinforcement learning; and ranking and selection, which can be used as the node selection policy in Monte Carlo tree search. We also review the literature on inventory management, which has been studied in both simulation optimization and AI.
Keywords: Simulation optimization, Artificial intelligence, Stochastic gradient estimation, Ranking and selection, Deep Learning, Reinforcement Learning
1. Introduction
Simulation is a powerful tool for analyzing complex stochastic systems found in manufacturing and supply chains, transportation, healthcare, call centers, finance, and many other fields [1], [2], [3]. Simulating a complex stochastic system is computationally expensive, and accurate system performance evaluation often requires a large number of simulation replications. Optimization of such complex stochastic systems may involve a large number of alternatives and decision variables. In recent decades, a stream of literature, e.g., [4], [5], [6], [7], has attempted to combine methods in simulation and optimization efficiently.
Inventory management is a central area in operations management. In a large e-commerce platform, there could be millions of stock keeping units (SKUs) stored in a complex network of warehouses. Simulation can be used to evaluate the cost of an inventory system. How to make reordering decisions for numerous SKUs to minimize the inventory costs is a complicated optimization problem.
Optimizing the reordering decisions is a Markov decision problem (MDP), which typically suffers from the curse of dimensionality in computation. To address this computational difficulty, many works focus on a simple policy such as base-stock and (s,S) policies that are only optimal under strict conditions that may not hold in practice, e.g., identically and independently distributed (i.i.d.) demands. In this case, artificial intelligence (AI) techniques have the potential to solve the original MDP. A new paradigm for solving the MDP is deep reinforcement learning (DRL), which parameterizes the reordering policy by an artificial neural network (ANN), where numerous parameters can be tuned through simulation optimization.
In this paper, we review two classic areas in simulation optimization: stochastic gradient estimation [8] and ranking and selection (R&S) [9]. We will discuss how these methods underpin modern AI techniques and the different focuses of simulation optimization and modern AI studies. The central issues addressed in AI have differed from those in other more traditional applications of simulation optimization, which will be discussed in detail in the following sections. We illustrate the application background using inventory management, a central problem in supply chain management, which has been studied in simulation [10], [11], [12], [13] and recently in AI [14], [15], [16], [17].
Prior to the recent buzz of ChapGPT, the most celebrated success story in AI was the stunning success of AlphaGo and its successors, developed using ANNs and Monte Carlo tree search (MCTS) [18], [19], [20], and trained by DRL to approximate value and policy functions. For an introduction to MCTS, see Fu [21], and for more information on ANNs, see Goodfellow et al. [22]. To a large extent, ANNs, MCTS, and DRL used in AlphaGo define modern AI.
ANNs are foundational to modern AI, and their training can be formulated as a simulation optimization problem, where stochastic gradient estimation is key for training them in deep learning and reinforcement learning [23], [24]. The upper confidence bound algorithm for trees (UCT) [25], derived from [26], [27] based on the multi-armed bandit (MAB) problem, is a popular node selection policy in MCTS [28], [29], [30], and is used in AlphaGo.
The rest of this paper is organized as follows. Section 2 introduces the general problem setting in simulation optimization. An overview of simulation optimization approaches and their connection to AI is provided in Section 3. Sections 4 and 5 introduce the two areas of stochastic gradient estimation and simulation budget allocation and their applications to AI. The last section concludes and discusses future directions. Most of the material of this review paper is taken from [31], including many parts verbatim and Figs. 2, Fig. 5, Fig. 6, Fig. 7, Fig. 8–9.
Fig. 2.
99% confidence intervals for the means of five alternatives.
Fig. 5.

Two tic-tac-toe board configurations to be considered.
Fig. 6.
Game tree for the first tic-tac-toe board configuration (assuming “greedy optimal” play), where we note that if the opponent’s moves were completely randomized, “O” would have an excellent chance of winning!.
Fig. 7.
Decision tree for the first tic-tac-toe board configuration ofFig. 5.
Fig. 8.
Decision tree for the first tic-tac-toe board configuration with non-optimal greedy moves.
Fig. 9.
Four steps in Monte Carlo tree search.
2. Problem setting
A generic simulation optimization problem can be written as
where is the feasible space of controllable parameters , and represents the randomness of the simulation model . In a simulation study, cannot be evaluated analytically but can be estimated by the average of simulation replications, i.e., , where , , represent the i.i.d. inputs of the simulation model. Given a parameter , running a simulation would typically be computationally expensive. The feasible space could be continuous or discrete. Continuous feasible spaces face computational challenges when the dimension of is high; similarly for discrete feasible spaces when the cardinality is large. Thus, the main challenge in solving the simulation optimization problem arises from the computational burden for evaluating the expected performance of the simulation model and searching for the optima of the expected performance.
We consider a periodic-review serial supply chain system containing intermediate echelons between a manufacturer and customers. An example with is shown in Fig. 1. We assume that the inventory replenishment at any stage takes the same amount of a fixed lead time denoted by . For echelon at period , we use , , and to denote on-hand inventory at the beginning, backlogs (accumulative unmet demand) at the beginning and ordering quantity, respectively. The quantity of goods that are shipped is
Actor makes decision for echelon . The update of the backlog and inventory for actor can be written as
Cost of echelon is given by
where coefficients of backlog costs and holding costs are given by and , respectively. The total costs of the system in period can be expressed by
Fig. 1.
A serial supply chain system with three intermediate echelons.
The randomness of the inventory system consists of customer demand . A simple inventory policy is the policy. Under an inventory control policy, when the inventory position (which includes inventory on hand plus that on order) falls below at an order decision point (discrete points in time in a periodic review setting and any point in time in a continuous review setting), then an order is placed for the amount that would bring the inventory position up to [32]. For the policy, the controllable vector (set of decision variables) is two-dimensional, i.e., . For DRL, the inventory policy can be represented by an ANN where there are numerous parameters, so the vector is high-dimensional. For the inventory system, the output of the simulation model is . How to optimize the inventory policy is transferred to optimize . Detailed description about the inventory system of the serial supply chain and more complicated inventory systems can be found in [15] and [16].
Next we turn to a recent application of simulation optimization, i.e., training an ANN in AI. Examples of output include a multidimensional vector for a classification problem that groups subjects into classes, or scalar output from a regression problem that estimates relationships among variables. The input could be an image represented by a vector of pixels or text encoded by a certain rule. The inputs are typically randomly drawn from a large dataset, e.g., ImageNet [33]. An ANN is comprised of many layers of neurons, the basic computational unit. ANN models usually contain randomness such as dropouts that randomly cut off the connection between two neurons and neuronal processing noises injected into the computation of neurons (see e.g., Peng et al. [24]). The computational complexity of ANNs arises from the composition of a huge number of neurons that are layered, as the output vector of the neurons in the preceding layer is the input vector for the neurons in the subsequent layer. From the discussion above, an ANN can be viewed as a stochastic simulation model parameterized by synaptic weights. For deep learning, training an ANN requires solving the following simulation optimization problem:
where is the loss function between the corresponding th output of the ANN, , and the th observation , is the vector of parameters. For regression, could be the quadratic loss function, and for classification, could be the cross-entropy loss function.
ANNs are also used in reinforcement learning to parameterize the value function and action policy. An MDP, which is used to model the reinforcement learning problem, can be represented as a 5-tuple , where and are state and action spaces, respectively; is the transition probability with being the current state and being the future state one step ahead; is the reward function; is the reward discount factor. One might aim to optimize a parameterized policy function , which is a distribution over the action space conditioned on the state, through interaction with the system. We denote the state and action at time by and , where and . The trajectory generated by the MDP can be thus defined as . Here we have , where is the initial state. The total reward of the trajectory is denoted as . In an episode, a complete simulation trajectory of the action and state is generated by the interaction between the agent and environment. For reinforcement learning, training an ANN requires solving the following simulation optimization problem:
| (1) |
3. Overview of simulation optimization approaches
With the problem setting and application background introduced, we provide a brief overview of sample average approximation (SAA), stochastic approximation (SA), and R&S for solving simulation optimization problems.
3.1. Overview of SAA and SA
The first two approaches are for solving simulation optimization problems with continuous controllable parameters. For example, for the single item inventory control problem, the two parameters can be optimized on the real space with constraints and . When the dimension of the parameter space is high, gradient-based optimization techniques are computationally preferable for solving generic nonlinear optimization problems. Deterministic gradient search can be applied to solve the nonlinear optimization via SAA as follows [34], [35]:
which closely resembles SA [36], [37]:
where , , are unbiased (pathwise) stochastic gradient estimates given independently and identically distributed (i.i.d.) inputs of the simulation model, is a gradient estimator for (e.g., ), and is the step size (learning rate). Comparisons between the performance of SAA and SA can be found in the literature, e.g., [38].
The SA approach with an unbiased stochastic gradient estimator (i.e., ) is referred to as the Robbins–Monro algorithm, and the SA with the (biased) finite-difference (FD) gradient estimator is called the Kiefer–Wolfowitz algorithm. The FD method for gradient estimation is always implementable without the need for information on the stochastic model. However, FD may have undesirable statistical properties, and its computational burden increases linearly with the dimension of the controllable parameters. The Kiefer–Wolfowitz algorithm requires tuning more parameters than the Robbins–Monro algorithm. These drawbacks make the Kiefer–Wolfowitz algorithm unfavorable for solving large-scale simulation optimization problems. For training an ANN, the Robbins–Monro algorithm and its variant are predominant. Recently, Hu et al. [39] provide a three-timescale stochastic approximation to optimize a quantile, based on two unbiased estimators for the gradient of the distribution function and density function, and Jiang et al. [40] develop a two-timescale stochastic approximation for optimizing a quantile in the context of DRL, based on the unbiased estimator for gradient of the distribution function. A weakness of gradient-based optimization methods is that they can only guarantee convergence to a local optimum. Some metaheuristics, such as simulation annealing, can guarantee convergence to a global optimum. In high-dimensional optimization problems such as training ANNs, convergence to a global optimum would often be computationally intractable, whereas convergence to a local optimum would usually be satisfactory enough.
Training an ANN typically involves the empirical minimization of a loss function by tuning the parameters. Factors such as the number of layers, the type of activation, and the layer sizes are chosen based on domain-specific applications [22]. However, SAA is generally not practical for training ANNs, because the input data may be huge (e.g., an image dataset) and the computational machinery of an ANN is heavy; hence, computing the summation of gradients of ANN over all data, i.e., , requires too much memory and computation to update the parameters. On the other hand, the update in SA requires much less memory and computational resources. However, to accelerate training the ANN, a mini-batch of inputs , , is put into a tensor, and a parallel computing platform can be leveraged to process the computation of ANN and the stochastic gradient estimate with the tensor as input. In practice, the entire input dataset is split into mini-batches with pieces of data, iterations of stochastic gradient descent using data in a mini-batch are run in an epoch, and random reshuffling of data is done at the end of an epoch before repeating the iterative process. The (global or local) convergence of SA is established under an assumption on step sizes (e.g., , ) and an assumption on the randomness of the stochastic gradient estimator [36]. The random reshuffling of data at the end of an epoch would generally make the randomness of the stochastic gradient estimator violate the assumptions required to ensure the convergence of SA. Tuning the step sizes in SA for training a large-scale ANN on a large dataset is usually a time-consuming and labor-intensive manual task. The step size is often kept constant in the beginning of the training and decreased after the loss decreases to a certain level in the later stages of the training. AdaGrad and Adam, variants of SA, introduce momentum and scaling factors for different coordinates to facilitate training the ANN [41], [42].
3.2. Overview of ranking and selection
We next turn our attention to R&S, which aims to solve discrete simulation optimization problems via efficient sampling. Specifically, our goal is to solve maximization problem , where the expectation of alternative is unknown but can be estimated by the average of independently and identically distributed simulated samples , . In a queueing network, scheduling servers at different queues to improve the performance is a combinatorial optimization problem with discrete controllable parameters. For discrete parameters where there is no information available on the local topology of a solution, finding the best alternative may require pairwise comparisons between different alternatives [9]. The problem of allocating computing resources to competing alternatives for finding the best alternative is a research question often called the R&S problem in the simulation literature, and it is closely related to the MAB problem in machine learning. In R&S, the primary goal has to do with correct selection of the best alternative, whereas traditional MAB work aims to maximize cumulative reward or minimize cumulative regret. Applications particularly suitable for methods in R&S include drug testing to guarantee a high probability that a lifesaving drug is effective. However, such methods may not be best suited for applications like designing a complex manufacturing system, where the possible number of alternative designs could be huge, running a simulation model would be expensive, and a managerial decision might need to be made quickly and adaptively to varying market conditions.
R&S techniques impact the efficiency of node selection in Monte Carlo tree search (MCTS). The successful Go-playing AI algorithm AlphaGo employs MCTS using the UCT node selection policy [18], where the number of possible states is of the order of , more than the number of atoms in the universe. MCTS provides a more effective way than traditional deterministic tree search techniques for estimating the state-action value function(s) via Monte Carlo simulation by following procedures of node selection and rollout. To improve the node selection policies in MCTS, techniques such as the Optimal Computing Budget Allocation (OCBA) and the Asymptotically Optimal Allocation Policy (AOAP) from R&S have been applied to MCTS, demonstrating better performance than the original UCT node selection policy [29], [30].
For simulation-based reinforcement learning algorithms to perform well, the state space should be explored sufficiently. Q-learning [43] also requires thorough exploration of the action space. Since pure greedy policies may lead to an insufficient exploration of the state-action space, -greedy, Boltzmann machines [43], and other MAB-based methods have been used to encourage more exploration. R&S-based methods have also been used to better balance exploration and exploitation [44], [45].
4. Stochastic gradient estimation
In this section, we focus on stochastic gradient estimation, the key element in implementing gradient-based simulation optimization methods such as SA and SAA introduced in the last section. For an introduction, see Fu [8]. Traditional application areas for stochastic gradient estimation include discrete event systems [46], financial engineering, and risk management [47], [48], [49], [50]. Fu [10] studies the stochastic gradient estimation for the two parameters in the inventory system. Recently, this topic has attracted attention in AI and machine learning; see a review paper by DeepMind [51]. In Section 4.1, we introduce classic problems and methods in stochastic gradient estimation, and we show how the methods can be applied to deep learning and reinforcement learning in Sections 4.2 and 4.3, respectively.
4.1. Classic problems and methods
Stochastic gradient estimation methods aim to estimate the gradient of the expected performance of a stochastic model . The infinitesimal perturbation analysis (IPA) or pathwise gradient is given by , and under appropriate conditions justifying the interchange of gradient and expectation operators, it is unbiased, i.e., satisfies
The required conditions for unbiasedness stem from applying the dominated convergence theorem, which for the IPA estimator generally translates to requiring that the performance function be Lipschitz continuous almost surely [52]. When is discontinuous, smoothed perturbation analysis (SPA) can be applied to derive an unbiased gradient estimator by conditioning on certain quantities to smooth out the discontinuities [53].
Suppose the density of has support , which is independent of the parameter . When the stochastic model depends on the parameter only through the input variables, i.e., , the gradient of the expected performance of a stochastic model can be given by
under certain conditions to justify the interchange of the gradient and expectation. The likelihood ratio (LR) gradient estimator is . Since the parameter does not directly appear in , the LR estimator does not require the continuity of to be unbiased. When the parameter appears directly in , the push-out LR could be applied, which requires an explicit change of variables to push out the parameter . Interested readers may refer to [54] for details of push-out LR.
When both IPA and LR are unbiased, as a rule-of-thumb, IPA has a smaller variance than LR [46], [48]. Cui et al. [55] provide sufficient conditions under which IPA has smaller variance, as well as a counterexample demonstrating that LR has a smaller variance when the sufficient condition is not satisfied. L’Ecuyer [56] provides a framework to unify IPA and LR, which is referred to as the combined derivative estimator in a recent review by Glasserman [57]. When is dependent on the parameter , Wang et al. [50] offer a technique to push out the parameter using a change of variables. Peng et al. [58], [59] consider an explicit stochastic model , where is a vector of functions with second-order differentiability and is a measurable function not necessarily continuous, and derive an unbiased generalized likelihood ratio (GLR) estimator in this setting. GLR generalizes LR by allowing to directly appear in the stochastic model and generalizes IPA by allowing discontinuities in the stochastic model. When the parameter does not directly appear in the stochastic model, GLR reduces to LR; when there is an explicit change of variables to push out the parameter, GLR coincides with the push-out LR, so GLR can be viewed as a generalization of push-out LR, which “implicitly” pushes out the parameter. See also [60] for some recent development. Because is only required to be a measurable function, which is very general, this flexibility allows GLR to handle many applications with discontinuities under a single umbrella.
Stochastic gradient estimation research for quantiles, conditional value-at-risk, and distortion risk measure is relatively recent [61], [62], [63], [64], [65], and is largely dependent on estimating the gradient of the distribution function [66], [67]. Stochastic gradient estimators for risk measures based on a single simulation run are generally biased. Much research attention has focused on analyzing the asymptotic properties of the estimators.
Traditional theoretical discussions in stochastic gradient estimation focus on the statistical properties of the estimator, such as unbiasedness and small variance. Less attention has been paid to the computational complexity, parallelization, and scalability of the estimator, which are critical for applications in AI. This may be due to the relatively lower-dimensional problems and the lack of well-accepted large-scale testbeds for analyzing the computational properties of estimators. In contrast, the AI community is built upon an open source ecosystem, where numerous ANN models, training algorithms, and datasets can be accessed online, some of which become benchmarks and testbeds of successful applications of AI such as computer vision (CV) and natural language processing (NLP) [68], [69], [70]. The testbed SimOpt, available at https://github.com/simopt-admin/simopt/wiki [71], [72], [73], [74], serves as an analogous valuable resource for the simulation optimization research community, with a diverse set of vetted examples, although relatively fewer large-scale industrial-level applications that are frequently found in AI testbeds.
4.2. Stochastic gradient estimation for deep learning
Training ANNs is central for deep learning, a prominent area in AI. Backpropagation (BP), an important gradient estimation technique embedded in most open source machine learning platforms for training ANNs, is shown by Peng et al. [24] to be pathwise equivalent to IPA, and backward propagation of errors can reduce computational complexity. Due to the nonlinearity of the activation function, tuning parameters in training ANNs is handled by SA algorithms. As there are often many parameters (the synaptic weights and biases to connect the neurons), the computational complexity in each descent step could be gigantic. In this regard, the BP method [75] provides an efficient way to compute the gradient at each step. It systematically applies the chain rule to propagate the perturbation effect of each parameter, by first making a forward pass to evaluate the function signals through the neuron layers, followed by a backward pass of the derivative components, or the so-called error signals, that ultimately allows an evaluation of the gradient with respect to all parameters simultaneously. Specifically, derivatives of all parameters can be decomposed into products of function signals and error signals, which results in reduced computational complexity relative to IPA without backward propagation of errors. Interested readers may refer to [24] for technical details.
Peng et al. [24] show that an LR-based (push-out) method can significantly improve the robustness of an ANN. The LR estimator is unbiased even if or is discontinuous. This provides substantially more flexibility in choosing and . For example, we may choose to be the zero-one loss function and to be a threshold function. Moreover, LR avoids the issue of vanishing gradient, which may arise in BP. However, the LR estimator is demonstrably simpler, as it does not need a backward recursion as in BP. The calculations of the LR-based estimator for parameters in each neuron are performed completely inside the neuron, which can be parallelized and results in substantial computational savings for higher-order derivatives. Interested readers may refer to [24] for technical details. By applying variance reduction techniques such as antithetic variates, Jiang et al. [76] have extended the LR method in [24], which was developed for a multi-layer perceptron (MLP), a basic ANN structure, to more complicated ANNs including CNNs, RNNs, graph neural networks (GNNs), and spiking neural networks (SNNs) with applications including CV, NLP, and social network analysis. ANNs trained by the LR-based method are more robust to perturbations in the input data than those trained by BP. Because the LR method only requires the local information of the model and the final loss for training ANN, it has a significant computational advantage in fine-tuning a module of a large ANN model.
4.3. Stochastic gradient estimation for reinforcement learning
Reinforcement learning, particularly DRL, has recently achieved many breakthroughs in AI. Policy gradient methods, an important class of DRL algorithms, solve problem (1) using stochastic gradient ascent. The key to implementing policy gradient methods is to estimate the gradient. Specifically, the gradient can be expressed as follows:
| (2) |
where is the trajectory of one episode in the MDP, and the interchange of the gradient and integral in the second equality can be justified by the dominated convergence theorem [77]. The term inside the expectation on the right-hand side of (2) is the LR estimator. We rewrite (2) as
where is the cost-to-go. Intuitively, the equality holds because the action determined by at time would not affect the returns before time but only affect the cost-to-go. The DRL algorithm that uses the gradient estimator above in the SA to optimize the parameter in ANN is called REINFORCE [78].
To improve data efficiency for the policy gradient method, we rewrite (2) as
To reduce the variance of gradient estimator, a baseline is subtracted from the cost-to-go :
This baseline independent of does not affect the expectation of the stochastic gradient estimator. The term can be viewed as a control variate that has zero expectation when is not controlled by and tends to be positively correlated with the stochastic gradient estimator, so it does not alter the expectation of the stochastic gradient estimator but can reduce its variance. For the actor-critic type of algorithms [43], the baseline is approximated by an ANN that is trained by the samples of the cost-to-go . The clip technique used in proximal policy optimization (PPO) can further reduce the variance of the policy gradient algorithm [69].
Jiang et al. [40] have developed policy gradient methods for quantile-based DRL by a two-time-scale SA algorithm. To improve the learning efficiency, Jiang et al. [40] use the bootstrap technique to allow multiple updates of parameters in one episode and truncate the trajectory of an episode into trajectories of different lengths, which are randomly picked by bootstrapping. Importance sampling and the clip technique are also applied in [40] to propose a quantile-based PPO (QPPO).
To apply DRL to the inventory management, the ordering quantity of the th echelon is treated as actor ’s action and a discrete action space for each actor is , where denotes the maximum ordering quantity in each period. The on-hand inventory, backlogs, downstream demand, and in-transit orders are included in the state of the MDP for optimizing the inventory ordering policy. When the action follows a policy modeled by a neural network , the output is parameterized by in the neural network. Rather than optimizing the two-dimensional parameter in the inventory policy, training the neural network requires optimizing the high-dimensional parameter , which could be performed by DRL. The costs and are used in the reward of the DRL. Wang et al. [79] and Liu et al. [15] apply an DRL and multi-agents DLR to optimize the expected cost of an ordering policy for a multi-echelon inventory system with and without information sharing in execution, respectively. Liu et al. [16] apply PPO to solve a large-scale inventory management problem of an e-commerce company with a complicated warehouse network. Zheng et al. [17] propose a dual-agent DLR algorithm to solve a joint dynamic pricing and inventory problem. Jiang et al. [40] apply QPPO to optimize the quantile of an ordering policy for a multi-echelon inventory system.
5. Ranking and selection
In this section, we provide a brief overview of R&S, an actively studied field in simulation optimization in the discrete input setting. For more information, readers can consult established reference volumes such as Bechhofer et al. [80], Chen and Lee [81], and Powell and Ryzhov [82]. Traditional applications of R&S include optimizing complex discrete event dynamic systems (DEDS) that are computationally intensive to simulate [81], and finding the most effective drugs, where the economic cost of each sample for testing the effectiveness of the drug is expensive [82]. Recent applications in AI can be found in Xiang et al. [83]. In Section 5.1, we introduce classical problems and methods in R&S, and we show how an AOAP method is adapted to be the selection policy of MCTS in Section 5.2.
5.1. Classical problems and methods
The goal of a classical R&S problem is to select the best alternative with the smallest (or largest) mean from a finite number of alternatives with unknown means . In the context of simulation, each alternative is typically a complex stochastic model (e.g., a queueing network), so the mean of the random output of the model (e.g., average system time) does not have an analytical form but can be estimated by simulation. The unknown means of alternatives can be estimated by independent random samples (simulation replications) following a joint distribution, i.e., , , where contains all the unknown parameters of the sampling distribution, including . The most common assumption in the literature of R&S is that the samples of different alternatives, i.e., , , follow independent normal distributions. In practice, the output of a stochastic model is often an average of many random variables (e.g., an average system time of 100 customers), so that the normal assumption of samples is justified by the central limit theorem. If the samples do not follow a normal distribution, then a macro-replication, which is the average of a batch of simulation replications, is used as a sample in the experiment; this macro-replication is approximately normally distributed. By the law of large numbers, sample means converge to means almost surely (a.s.), i.e., a.s., as , . Thus, as the simulation budget goes to infinity, the best alternative could be eventually selected. However, simulation replications are typically expensive, which often leads to a very limited simulation budget in practice.
Fig. 2 presents the corresponding 99% confidence intervals for the means of five alternatives. In this example, it is highly unlikely for alternatives 1, 4, and 5 to be the best alternative (with the lowest mean). Therefore, it seems unnecessary to spend a significant number of simulation replications on accurately learning the means of alternatives 1, 4, and 5 rather than focusing on learning the means of the other two more promising alternatives to distinguish which one is actually the best. This example motivates the study of a central issue in the research of R&S, i.e., how to allocate simulation replications intelligently to improve sampling efficiency. Intuitively, allocating more simulation replications to the alternatives with smaller means and higher variances seems to be reasonable for efficiently selecting the best alternative. In the literature, this intuition is summarized as the “mean-variance” or “exploitation-exploration” tradeoff. To formalize the study on the issue, we need to introduce a metric to measure sampling efficiency. In R&S, a popular metric is the probability of correct selection (PCS), i.e., , where is the selection decision made after allocating a total of simulation replications, is the information set obtained after allocating simulation replications, and is the zero-one loss after allocating simulation replications with . Discussion of different metrics can be found in Eckman and Henderson [84]. Notice that for the metric commonly used in R&S, the loss only occurs at the end after allocating all simulation replications.
The MAB problem is closely related to R&S [85], [86], [87]. The goal of a classic MAB problem is similar to that of R&S, i.e., to play most often the best arm from a finite number of arms with unknown expected rewards (means). A major difference from R&S lies in the objective, which for traditional MAB is the total expected regret (or reward) accumulated each time an arm is pulled, although more recently MAB has also considered a setting termed best-arm identification (BAI) or the pure exploration problem, where a correct selection criterion is also used. A good balance of exploitation and exploration is also key to intelligently sampling each alternative in the MAB problem. Thompson sampling (TS) [88] is a popular sampling procedure in the traditional MAB setting, but the performance of TS is notoriously poor for R&S or BAI problem [87]. Recently, Russo [89] proposed a top-two Thompson sampling (TTTS) procedure that performs well for the R&S problem. Compared to TS, TTTS tends to explore more when allocating samples. Shi et al. [90] extend TTTS to select contextually dependent top- alternatives.
In R&S, numerous existing sampling procedures can be basically categorized into two branches developed from two distinct perspectives, which can be viewed as duals of each other. An earlier branch, studied under an indifference zone (IZ) paradigm [91], [92], aims to guarantee a predetermined PCS level for parameters satisfying an IZ assumption, i.e., with . Since the sampling procedures derived in the IZ paradigm need to guarantee a PCS level for the worst-case scenario, called the slippage configuration [93], in practice these procedures tend to be overly conservative in that they often require allocating more simulation replications than necessary for guaranteeing a specified PCS level. A recent theoretical advance by Fan et al. [94] eliminates the need for the IZ assumption in guaranteeing PCS, although the drawback of conservatism still persists.
Rather than focusing on the worst-case scenario, Chen [95] (see also Chen et al. [96]) developed an alternative approach called the optimal computing budget allocation (OCBA) method. OCBA solves a static optimization problem, i.e.,
with being the number of simulation replications allocated to alternative . With several approximations, an asymptotically optimal sampling ratio with an analytical formula can be obtained, given the unknown parameter [97], [98]. Subsequently, an estimate for is incorporated into the OCBA formula. This is done using initial samples that are usually equally distributed to each alternative in the first stage. As a result, the remaining samples can be allocated based on the sampling ratio suggested by OCBA, thereby enhancing sampling efficiency. Here the sampling efficiency should not be defined as the PCS, given a specific parameter . For a given parameter, we can directly select the best alternative by sorting means which are a part of without running simulations. Therefore, a reasonable performance metric could be an integrated PCS (IPCS) over a weighting prior measure on the parameter space , i.e., . Under a Bayesian framework, a dynamic allocation and selection (A&S) policy to optimize IPCS is defined in Peng et al. [99], and a Bellman equation for solving the optimal A&S policy is rigorously established in Peng et al. [100]. Well-known Bayesian sequential sampling procedures include expected value of improvement (EVI) in Chick et al. [101], knowledge gradient (KG) in Frazier et al. [102], and expected improvement (EI) in Ryzhov [103]. KG and EI sequentially allocate a sample to optimize myopic surrogate criteria.
The dynamic decisions in the R&S problem can be formulated as an A&S policy [99]. The allocation policy is a sequence of mappings that sequentially allocates each sample to an alternative based on collected information, and the selection policy makes a final decision to select the best alternative after exhausting all simulation replications. From Peng et al. [100], the optimal A&S policy satisfies a Bellman equation. The commonly assumed selection policy chooses the alternative with the largest posterior mean, even though it is not the optimal selection policy.
Peng et al. [100] propose a dynamic sampling procedure derived in an approximate dynamic programming (ADP) paradigm. The idea of ADP is to find a suitable value function approximation (VFA) for the Bellman equation. A simple treatment is to fix the selection policy to select the alternative with the largest posterior mean, and look one step ahead for allocating the next replication. Subsequently, VFA becomes an integral of the standard normal density over a region encompassed by hyperplanes. The VFA can be further simplified to be an analytical form. AOAP has an analytical form that reflects a similar (posterior) mean-variance tradeoff as the OCBA formula. Not only is AOAP consistent but it also sequentially achieves the asymptotically optimal sampling ratio.
Beyond the classic R&S problem, recent advances have been made in subset selection [104], [105], [106], feasibility identification [107], [108], quantile-based simulation sample allocation [109], [110], context-dependent simulation sample allocation [90], [111], [112], [113], and large-scale parallel computing simulation sampling allocation [114], [115], [116], [117]. Many of these recent advances have not yet been applied to AI, where they have the potential for improving computational efficiency.
The fixed-budget R&S problem as MDP can be characterized by the following Bellman equation: for ,
where is the optimal allocation policy at (t+1)-th step, is the state-action-value function, which represents the reward value of choosing alternative , and is determined by and newly allocated observation. Zhou et al. [118] propose to solve the MDP via a rollout method depicted in Fig. 3, where the state-action-value function is estimated by an average of simulated sample paths following a rollout policy with the only reward received at the end, i.e., 0 if the best alternative is incorrectly selected and 1 if the best is selected correctly.
Fig. 3.
The decisions in the rollout process.
The rollout policy is much slower compared with the traditional R&S method. Zhou et al. [118] propose to use an offline-trained ANN for learning the value functions estimated by the rollout policy and employing the ANN to solve the SDP online. In Fig. 4, a neural network is trained to specify the allocation policy and base policy for rollout, and another neural network is trained to evaluate the PCS for implementing the stopping rule. In the pre-training procedure called AlphaRank, the input, output, and training objectives of the ANN are defined first; then labeled data is selected in the training process; finally reasonable evaluation criteria is defined to quantify the performance of the ANN model after training. This research demonstrates that not only simulation optimization contributes to recent advances in AI but also recent AI advancements can inspire novel approaches in simulation optimization.
Fig. 4.
NN training and evaluating architecture with the number of alternatives = 2.
5.2. Ranking and selection for Monte Carlo tree search
As mentioned earlier, AlphaGo employs MCTS, and simulation sample allocation is important for node selection in MCTS. We illustrate the MCTS for a board game using the simple example of tic-tac-toe, which is a well-known two-player game on a 3 3 grid. Players alternate turns, with the “X” player going first, aiming to get 3 marks in a row. Assuming the primary objective (for both players) is to win and the secondary objective is to avoid losing, the following “greedy” policy is optimal (for both players):
If a win move is available, then take it; else if a block move is available, then take it.
However, there remain numerous situations where neither a win move nor a block move is available. In traditional game tree search, these would be enumerated. We instead illustrate the MCTS approach, which samples the opposing moves rather than enumerating them.
When going first (as X), there are three unique first moves to consider – corner, side, and middle. If going second (as O), the possible moves depend of course on what X has selected. If the middle was chosen, then there are two unique moves available (corner or non-corner side); if a side (corner or non-corner) was chosen, then there are five unique moves available. Despite the game having a relatively small number of outcomes compared to most other games, such as checkers, enumerating them all for illustration purposes can be quite messy, so we further simplify by examining two different game situations that already have preliminary moves.
Assume henceforth that we are the “O” player. We begin by illustrating with two games that have already seen three moves, two by our opponent and one by us, so it is our turn to make a move. The two different board configurations are shown in Fig. 5. For the first game, by symmetry there are just two unique moves for us to consider: corner or (non-corner) side. In this particular situation, following the optimal policy above leads to an easy decision: corner move leads to a loss, and non-corner move leads to a draw; there is a unique “sample path” in both cases, thus eliminating the need to simulate. The trivial game tree and decision tree are provided in Figs. 6 and 7, respectively, with the game tree that allows non-optimal moves shown in Fig. 8. A more complicated game configuration can be found in Fu [21], and an online interactive Java demo of MCTS can be found at https://terpconnect.umd.edu/∼mfu/demo.
We view each configuration of a board game as a node and each move as an edge in MCTS. In general, MCTS consists of four steps: selection, expansion, simulation, and back-propagation, which are depicted in Fig. 9. The functions of the four steps are described below:
-
•
Selection: the MCTS algorithm traverses the current tree from the root node following a tree policy. The policy employs an evaluation function to optimally select nodes with the highest estimated value.
-
•
Expansion: a new child node is added to the tree to that node which was optimally reached during the selection step.
-
•
Simulation: a simulation is performed following a default policy until a result or predefined state is achieved.
-
•
Back-propagation: after determining the value of the newly added node, the remaining tree must be updated from the new node to the root node. The number of simulations stored in each node is incremented and if the new node’s simulation results in a win, then the number of wins is also incremented (similarly for a loss or draw).
Traditionally, the UCT algorithm in MAB has been used as the tree selection policy to balance the exploration and exploitation [25], whereas R&S procedures have been applied to the tree selection to further improve performance [29], [119]. In AlphaGo Zero, the policy network can assist the selection of actions. Specifically, to determine the next action at state in the th rollout, UCT selects the action with the highest upper confidence bound as
where is a positive exogenous parameter that balances the exploration and exploitation during node selection, and and are the sample mean and the number of visit times of the state-action pair of the state-action value after rollouts, respectively. After all rollouts are conducted, UCT determines the final action under the root state as the action with the highest number of visit times, i.e.,
An application of an AlphaGo-like MCTS to adaptive bitrate control in video delivery can be found in [120]. Applications of MCTS to inventory management and computational finance can be found in [121] and [122], respectively.
Liu et al. [30] propose Asymptotically Optimal Allocation Policy for Trees (AOAT) by extending AOAP to multi-stage node selection problem in MCTS, which utilizes sample means, variances, and the number of node visits to balance the exploration of nodes with large variances and the exploitation of nodes with large state-action values. The value network could provide prior knowledge for AOAT. A variant of AOAT called AOAT-Pi includes the policy network for determining the final action.
6. Conclusions
We reviewed simulation optimization methods closely related to AI, focusing on two areas: stochastic gradient estimation and R&S. We discussed some classic techniques in AI from the perspective of simulation optimization and introduced new methods for advancing modern AI techniques. The differences between traditional simulation optimization applications and new applications in modern AI were also discussed. Traditional simulation optimization research focuses on statistical properties of stochastic estimators (e.g., for the gradient) and asymptotic analyses of optimization algorithms, whereas the computational complexity, parallelization, and scalability of the estimator and finite-sample behavior of the optimization algorithm are deemed as more critical for AI applications. By bridging the gap, advances in simulation optimization can be used to improve training speed, robustness, and resilience of decision making in AI (see [76] and [40]).
Beyond AlphaRank in [118] for solving classical R&S problems, future research direction includes extending the AlphaRank framework to subset selection, feasibility identification, quantile-based simulation sample allocation, and context-dependent simulation sample allocation. The training of an AI agent is largely a mean-based optimization, whereas empirical behavioral studies have demonstrated that human decision making under uncertainty is inconsistent with expected utility models, which are not able to adequately account for risk. Quantile-based DRL proposed by [40] can better capture the tail risk. Recent research incorporating risk includes Glynn et al. [61], who consider a distortion risk measure that is a Lebesgue-Stieltjes integral of quantiles. Integrating risk measures into the training of AI agents would result in better aligned decisions; see also [123].
Declaration of competing interest
The authors declare that they have no conflicts of interest in this work.
Acknowledgments
This work was supported in part by the National Natural Science Foundation of China (72250065 and 72022001), by the U.S. Air Force Office of Scientific Research (AFOSR) (FA95502010211), and by Xiangjiang Laboratory Key Project (23XJ02004). This work was also supported in part by Wuhan East Lake High-Tech Development Zone (also known as the Optics Valley of China, or OVC) National Comprehensive Experimental Base for Governance of Intelligent Society.
Biographies
Yijie Peng is currently an associate professor of the Department of Management Science and Information Systems in Guanghua School of Management at Peking University (PKU). He received his Ph.D. from the Department of Management Science at Fudan University and his BS degree from the School of Mathematics at Wuhan University. Many of his publications appear in high-quality journals including Operations Research, INFORMS Journal on Computing, and IEEE Transactions on Automatic Control. He is awarded the 2019 Outstanding Simulation Publication Award of INFORMS Simulation Society. He serves as an associate editor for Asia-Pacific Journal of Operational Research and IEEE Control Systems Society Conference Editorial Board. His research interests include stochastic modeling and analysis, simulation optimization, machine learning, and healthcare.
Chun-Hung Chen received his Ph.D. degree from Harvard University in 1994. He is currently a professor at George Mason University. Dr. Chen was an assistant professor at the University of Pennsylvania before joining GMU. Sponsored by NSF, NIH, DOE, NASA, FAA, and AFOSR in the U.S., he has worked on the development of very efficient methodology for simulation-based decision making and its applications. Dr. Chen has served on several editorial boards, such as IEEE Transactions on Automatic Control, IEEE Transactions on Automation Science and Engineering, IIE Transactions, Asia-Pacific Journal of Operational Research, Journal of Simulation Modeling Practice and Theory, International Journal of Simulation and Process Modeling, and Journal of Traffic and Transportation Engineering. Dr. Chen is an author of the popular book Stochastic Simulation Optimization: An Optimal Computing Budget Allocation. He is an IEEE Fellow.
Michael C. Fu holds the Smith Chair of Management Science at the Robert H. Smith School of Business, with a joint appointment in the Institute for Systems Research, at the University of Maryland. He is the co-author of the books, Conditional Monte Carlo: Gradient Estimation and Optimization Applications, which received the INFORMS Simulation Society’s 1998 Outstanding Publication Award, and Simulation-Based Algorithms for Markov Decision Processes, and editor/co-editor of four volumes: Perspectives in Operations Research, Advances in Mathematical Finance, Encyclopedia of Operations Research and Management Science (3rd edition), and Handbook of Simulation Optimization. He served as the Operations Research NSF Program Director (2010–2012 & 2015) and as General Co-Chair for the 2020 INFORMS National Meeting. Recent awards include the INFORMS Simulation Society’s Distinguished Service Award (2019), the INFORMS Saul Gass Expository Writing Award (2021), and the INFORMS Kimball Medal (2022). He is a Fellow of IEEE and INFORMS.
Footnotes
Peer review under the responsibility of Editorial Board of Fundamental Research.
References
- 1.Asmussen S., Glynn P.W. Springer; New York: 2007. Stochastic Simulation: Algorithms and Analysis. [Google Scholar]
- 2.Law A.M., Kelton W.D. McGraw-Hill, Boston, MA; 2000. Simulation Modeling and Analysis. [Google Scholar]
- 3.Rubinstein R.Y., Kroese D.P. John Wiley & Sons; 2016. Simulation and the Monte Carlo Method. [Google Scholar]
- 4.Chen C.-H., Fu M.C., Shi L. State-of-the-art Decision-Making Tools in the Information-Intensive Age. INFORMS TutORials; 2008. Simulation and optimization; pp. 247–260. [Google Scholar]; https://doi.org/10.1287/educ.1080.0050.
- 5.Fu M.C. Optimization for simulation: Theory vs. practice (Feature Article) INFORMS J. Comput. 2002;14(3):192–215. [Google Scholar]
- 6.Fu M.C. Springer; New York: 2015. Handbook of Simulation Optimization. [Google Scholar]
- 7.Pasupathy R., Ghosh S. Simulation optimization: A concise overview and implementation guide. INFORMS TutORials Oper. Res. 2013:122–150. [Google Scholar]
- 8.Fu M.C. In: Handbook of Simulation Optimization. Fu M.C., editor. Springer; New York: 2015. Stochastic gradient estimation; pp. 105–147. [Google Scholar]
- 9.Hong L.J., Fan W., Luo J. Review on ranking and selection: A new perspective. Front. Eng. Manag. 2021;8(3):321–343. [Google Scholar]
- 10.Fu M.C. Sample path derivatives for (s, S) inventory systems. Oper. Res. 1994;42(2):351–364. [Google Scholar]
- 11.Glasserman P., Tayur S. Sensitivity analysis for base-stock levels in multiechelon production-inventor systems. Manag. Sci. 1995;41(2):263–281. [Google Scholar]
- 12.Bashyam S., Fu M.C. Optimization of (s, S) inventory systems with random lead times and a service level constraint. Manag. Sci. 1998;44(12-NaN-2):S243–S256. [Google Scholar]
- 13.Kunnumkai S., Topaloglu H. Using stochastic approximation methods to compute optimal base-stock levels in inventory control problems. Oper. Res. 2008;56(3):646–664. [Google Scholar]
- 14.Gijsbrechts J., Boute R., Van Mieghem J., et al. Can deep reinforcement learning improve inventory management? Performance on lost sales, dual sourcing and multi-echelon problems. Manuf. Serv. Oper. Manag. 2021;24(3):1261–1885. [Google Scholar]
- 15.Liu X., Hu M., Peng Y., et al. Multi-agent deep reinforcement learning for multi-echelon inventory management. Prod. Oper. Manag. 2024 doi: 10.1177/10591478241305863. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16.Liu X., Alexopolous C., Peng Y. A simulation-driven machine learning framework for large-scale inventory management. SSRN Electron. J. 2024 [Google Scholar]; Manuscript No. 4490327
- 17.Y. Zheng, Z. Li, P. Jiang, et al. Dual-agent deep reinforcement learning for dynamic pricing and replenishment, Preprint at arXiv: 2410.21109 (2024).
- 18.Silver D., Huang A., Maddison C.J., et al. Mastering the game of Go with deep neural networks and tree search. Nature. 2016;529(7587):484–489. doi: 10.1038/nature16961. [DOI] [PubMed] [Google Scholar]
- 19.Silver D., Hubert T., Schrittwieser J., et al. A general reinforcement learning algorithm that masters chess, Shogi, and Go through self-play. Science. 2018;362:1140–1144. doi: 10.1126/science.aar6404. [DOI] [PubMed] [Google Scholar]
- 20.Silver D., Schrittwieser J., Simonyan K., et al. Mastering the game of Go without human knowledge. Nature. 2017;550:354–359. doi: 10.1038/nature24270. [DOI] [PubMed] [Google Scholar]
- 21.Fu M.C. Simulation-based algorithms for Markov decision processes: Monte Carlo tree search from AlphaGo to AlphaZero. Asia-Pacific J. Oper. Res. 2019;36(06) [Google Scholar]
- 22.Goodfellow I., Bengio Y., Courville A. MIT Press; 2016. Deep Learning. [Google Scholar]; http://www.deeplearningbook.org.
- 23.Haykin S. Pearson Education; New Jersey: 2009. Neural Networks and Learning Machines, 3rd Ed. [Google Scholar]
- 24.Peng Y., Xiao L., Heidergott B., et al. A new likelihood ratio method for training artificial neural networks. INFORMS J. Comput. 2022;83(403):834–841. [Google Scholar]
- 25.Kocsis L., Szepesvári C. Machine Learning: ECML 2006: 17th European Conference on Machine Learning Berlin, Germany, September 18-22, 2006 Proceedings 17. Springer; 2006. Bandit based Monte-Carlo planning; pp. 282–293. [Google Scholar]
- 26.Auer P., Cesa-Bianchi N., Fischer P. Finite-time analysis of the multiarmed bandit problem. Mach. Learn. 2002;47(2):235–256. [Google Scholar]
- 27.Chang H.S., Fu M.C., Hu J., et al. An adaptive sampling algorithm for solving Markov decision processes. Oper. Res. 2005;53(1):126–139. [Google Scholar]
- 28.Coulom R. In: Computers and Games. CG 2006. Lecture Notes in Computer Science; Proceedings Computers and Games 2006. van den Herik H.J., Ciancarini P., Donkers H.H.L.M., editors. volume 4630. Springer; Berlin: 2006. Efficient selectivity and backup operators in Monte-Carlo tree search; pp. 72–83. [Google Scholar]
- 29.Li Y., Fu M.C., Xu J. An optimal computing budget allocation tree policy for Monte Carlo tree search. IEEE Trans. Autom. Control. 2022;67(6):2685–2699. [Google Scholar]
- 30.Liu X., Peng Y., Zhang G., et al. An efficient node selection policy for Monte Carlo tree search with neural networks. INFORMS J. Comput. 2024 [Google Scholar]; preprint at arXiv: 2204.12043
- 31.Peng Y., Chen C.-H., Fu M.C. Simulation optimization in the new era of AI. INFORMS TutORials Oper. Res. 2023:82–108. [Google Scholar]
- 32.Porteus E.L. Stanford University Press; Stanford: 2002. Foundations of Stochastic Inventory Theory. [Google Scholar]
- 33.P. Goyal, P. Dollár, R. Girshick, et al. Accurate, large minibatch SGD: Training ImageNet in 1 hour, arXiv preprint arXiv:1706.02677 (2017).
- 34.Kim S., Pasupathy R., Henderson S.G. A guide to sample average approximation. Chapter 8 in Handbook of Simulation Optimization. 2015:207–243. [Google Scholar]
- 35.Shapiro A., Dentcheva D., Ruszczynski A. SIAM; 2021. Lectures on Stochastic Programming: Modeling and Theory. [Google Scholar]
- 36.Kushner H.J., Clark D.S. Springer Science & Business Media; 2012. Stochastic Approximation Methods for Constrained and Unconstrained Systems. [Google Scholar]
- 37.Newton D., Yousefian F., Pasupathy R. Stochastic gradient descent: Recent trends. INFORMS TutORials Oper. Res. 2018:193–220. [Google Scholar]
- 38.Lam H., Jiang G., Fu M.C. Proceedings of the 2018 Winter Simulation Conference. IEEE; 2018. On efficiencies of stochastic optimization procedures under importance sampling; pp. 1862–1873. [Google Scholar]
- 39.Hu J., Peng Y., Zhang G., et al. A stochastic approximation method for simulation-based quantile optimization. INFORMS J. Comput. 2022;34(5):2889–2907. [Google Scholar]
- 40.J. Jiang, J. Hu, Y. Peng, Quantile-based deep reinforcement learning using two-timescale policy gradient algorithms, arXiv preprint arXiv:2305.07248(2023).
- 41.D.P. Kingma, J. Ba, Adam: a method for stochastic optimization, arXiv preprint arXiv:1412.6980 (2014).
- 42.Lydia A., Francis S. AdaGrad - an optimizer for stochastic gradient descent. Int. J. Inf. Comput.Sci. 2019;6(5):566–568. [Google Scholar]
- 43.Bertsekas D., Tsitsiklis J.N. Athena Scientific; 1996. Neuro-Dynamic Programming. [Google Scholar]
- 44.Jia Q.-S. Efficient computing budget allocation for simulation-based policy improvement. IEEE Trans. Autom. Sci.Eng. 2012;9(2):342–352. [Google Scholar]
- 45.Zhu Y., Dong J., Lam H. Efficient uncertainty quantification and exploration for reinforcement learning. Oper. Res. 2024;72(4):1689–1709. [Google Scholar]
- 46.Cassandras C.G., Lafortune S. Springer Nature; Switzerland: 2021. Introduction to Discrete Event Systems. [Google Scholar]
- 47.Chen N., Liu Y. American option sensitivities estimation via a generalized infinitesimal perturbation analysis approach. Oper. Res. 2014;62(3):616–632. [Google Scholar]
- 48.Glasserman P. Springer; 2004. Monte Carlo Methods in Financial Engineering. [Google Scholar]
- 49.Hong L.J., Juneja S., Luo J. Estimating sensitivities of portfolio credit risk using Monte Carlo. INFORMS J. Comput. 2014;26(4):848–865. [Google Scholar]
- 50.Wang Y., Fu M.C., Marcus S.I. A new stochastic derivative estimator for discontinuous payoff functions with application to financial derivatives. Oper. Res. 2012;60(2):447–460. [Google Scholar]
- 51.Mohamed S., Rosca M., Figurnov M., et al. Monte Carlo gradient estimation in machine learning. J. Mach. Learn. Res. 2020;21(132):1–62. [Google Scholar]
- 52.Glasserman P. Kluwer Academic Publishers, Boston; 1991. Gradient Estimation via Perturbation Analysis. [Google Scholar]
- 53.Fu M.C., Hu J.-Q. Kluwer Academic Publishers, Boston; 1997. Conditional Monte Carlo: Gradient Estimation and Optimization Applications. [Google Scholar]
- 54.Rubinstein R.Y., Shapiro A. volume 346. Wiley; New York: 1993. Discrete Event Systems: Sensitivity Analysis and Stochastic Optimization by the Score Function Method. [Google Scholar]
- 55.Cui Z., Fu M.C., Hu J.-Q., et al. On the variance of single-run unbiased stochastic derivative estimators. INFORMS J. Comput. 2020;32(2):390–407. [Google Scholar]
- 56.L’Ecuyer P. A unified view of the IPA, SF, and LR gradient estimation techniques. Manag. Sci. 1990;36(11):1364–1383. [Google Scholar]
- 57.P. Glasserman, Combined Derivative Estimators. In: Botev, Z., Keller, A., Lemieux, C., Tuffin, B. (eds) Advances in Modeling and Simulation. Springer, Cham, 2022, 10.1007/978-3-031-10193-9_10 [DOI]
- 58.Peng Y., Fu M.C., Hu J.-Q., et al. A new unbiased stochastic derivative estimator for discontinuous sample performances with structural parameters. Oper. Res. 2018;66(2):487–499. [Google Scholar]
- 59.Peng Y., Fu M., Hu J., et al. Generalized likelihood ratio method for stochastic models with uniform random numbers as inputs. Eur. J. Oper. Res. 2025;321(2):493–502. [Google Scholar]
- 60.Ren X., Fu M. Proceedings of the 2024 Winter Simulation Conference. IEEE; 2024. Generalizing the generalized likelihood ratio method through a new push-out Leibniz integration approach. [Google Scholar]
- 61.Glynn P.W., Peng Y., Fu M.C., et al. Computing sensitivities for distortion risk measures. INFORMS J. Comput. 2021;33(4):1520–1532. [Google Scholar]
- 62.Heidergott B., Volk-Makarewicz W. A measure-valued differentiation approach to sensitivity analysis of quantiles. Math. Oper. Res. 2016;41(1):293–317. [Google Scholar]
- 63.Hong L.J. Estimating quantile sensitivities. Oper. Res. 2009;57(1):118–130. [Google Scholar]
- 64.Hong L.J., Liu G. Simulating sensitivities of conditional value at risk. Manag. Sci. 2009;55(2):281–293. [Google Scholar]
- 65.Jiang G., Fu M.C. Technical note — On estimating quantile sensitivities via infinitesimal perturbation analysis. Oper. Res. 2015;63(2):435–441. [Google Scholar]
- 66.Liu G., Hong L.J. Kernel estimation of the Greeks for options with discontinuous payoffs. Oper. Res. 2011;59(1):96–108. [Google Scholar]
- 67.Peng Y., Fu M.C., Heidergott B., et al. Maximum likelihood estimation by Monte Carlo simulation: towards data-driven stochastic modeling. Oper. Res. 2020;68(6):1896–1912. [Google Scholar]
- 68.Deng J., Dong W., Socher R., et al. 2009 IEEE Conference on Computer Vision and Pattern Recognition. IEEE; 2009. ImageNet: A large-scale hierarchical image database; pp. 248–255. [Google Scholar]
- 69.J. Schulman, F. Wolski, P. Dhariwal, et al. Proximal policy optimization algorithms, arXiv preprint arXiv:1707.06347 (2017).
- 70.Vaswani A., Shazeer N., Parmar N., et al. Attention is all you need. Adv. Neural Inf. Process. Syst. 2017;30:6000–6010. [Google Scholar]
- 71.Dong N.A., Eckman D.J., Zhao X., et al. Proceedings of the 2017 Winter Simulation Conference. IEEE; 2017. Empirically comparing the finite-time performance of simulation-optimization algorithms; pp. 2206–2217. [Google Scholar]
- 72.Eckman D.J., Henderson S.G., Pasupathy R. Proceedings of the 2019 Winter Simulation Conference. IEEE; 2019. Redesigning a testbed of simulation-optimization problems and solvers for experimental comparisons; pp. 3457–3467. [Google Scholar]
- 73.Pasupathy R., Henderson S.G. Proceedings of the 2006 Winter Simulation Conference. IEEE; 2006. A testbed of simulation-optimization problems; pp. 255–263. [Google Scholar]
- 74.Pasupathy R., Henderson S.G. Proceedings of the 2011 Winter Simulation Conference. IEEE; 2011. SimOpt: A library of simulation optimization problems; pp. 4075–4085. [Google Scholar]
- 75.Rumerlhar D.E., Hinton G.E., Williams R.J. Learning representation by back-propagating errors. Nature. 1986;323:533–536. [Google Scholar]
- 76.Jiang J., Zhang Z., Xu C., et al. One forward is enough for neural network training via likelihood ratio method. International Conference on Representations Learning (ICLR) 2024 [Google Scholar]
- 77.Rudin W. McGraw-Hill, New York; 1964. Principles of Mathematical Analysis. [Google Scholar]
- 78.Sutton R., Barto A. The MIT Press; Massachusetts: 2014. Reinforcement Learning: An Introduction. [Google Scholar]
- 79.Wang P.Y., Qinghao, Yang Y. Solving inventory management problems through deep reinforcement learning. J. Syst. Sci. Syst. Eng. 2022;31(8):677–689. [Google Scholar]
- 80.Bechhofer R.E., Santner T.J., Goldsman D.M. New York: John Wiley and Sons; 1995. Design and Analysis for Statistical Selection, Screening, and Multiple Comparisons. [Google Scholar]
- 81.Chen C.-H., Lee L.H. World Scientific Publishing Company; Toh Tuck Link, Singapore: 2011. Stochastic Simulation Optimization: An Optimal Computing Budget Allocation. [Google Scholar]
- 82.Powell W.B., Ryzhov I.O. Chapter 4 in Optimal Learning. New York: John Wiley and Sons; 2012. Ranking and selection; pp. 71–88. [Google Scholar]
- 83.Xiang H., Lin J., Chen C.-h., et al. Asymptotic meta learning for cross validation of models for financial data. IEEE Intell. Syst. 2020;35(2):16–24. [Google Scholar]
- 84.Eckman D.J., Henderson S.G. Fixed-confidence, fixed-tolerance guarantees for ranking-and-selection procedures. ACM Trans. Model. Comput.Simul. (TOMACS) 2021;31(2):1–33. [Google Scholar]
- 85.Lai T.L. Adaptive treatment allocation and the multi-armed bandit problem. Ann. Stat. 1987;15(3):1091–1114. [Google Scholar]
- 86.Robbins H. Some aspects of the sequential design of experiments. Bull. Am. Math. Soc. 1952;58(5):527–535. [Google Scholar]
- 87.Russo D.J., Van Roy B., Kazerouni A., et al. A tutorial on Thompson sampling. Found. Trends Mach. Learn. 2018;11(1):1–96. [Google Scholar]
- 88.Thompson W.R. On the likelihood that one unknown probability exceeds another in view of the evidence of two samples. Biometrika. 1933;25(3-4):285–294. [Google Scholar]
- 89.Russo D. Simple Bayesian algorithms for best-arm identification. Oper. Res. 2020;68(6):1625–1647. [Google Scholar]
- 90.X. Shi, Y. Peng, G. Zhang, Top-two Thompson sampling for contextual top-mc selection problems, arXiv: 2306.17704 (2024).
- 91.Goldsman D., Nelson B.L. Chapter 8 in Handbook of Simulation: Principles, Methodology, Advances, Applications, and Practice, ed. J. Banks. New York: John Wiley and Sons; 1998. Comparing systems via simulation. [Google Scholar]
- 92.Kim S.-H., Nelson B.L. Chapter 17 in Handbooks in Operations Research and Management Science: Simulation. volume 13. Elsevier; 2006. Selecting the best system; pp. 501–534. [Google Scholar]
- 93.Branke J., Chick S.E., Schmidt C. Selecting a selection procedure. Manag. Sci. 2007;53(12):1916–1932. [Google Scholar]
- 94.Fan W., Hong L.J., Nelson B.L. Indifference-zone-free selection of the best. Oper. Res. 2016;64(6):1499–1514. [Google Scholar]
- 95.Chen C.-H. A lower bound for the correct subset-selection probability and its application to discrete-event system simulations. IEEE Trans. Autom. Control. 1996;41(8):1227–1231. [Google Scholar]
- 96.Chen C.-H., Lin J., Yücesan E., et al. Simulation budget allocation for further enhancing the efficiency of ordinal optimization. Discrete Event Dyn. Syst. 2000;10(3):251–270. [Google Scholar]
- 97.Glynn P.W., Juneja S. Proceedings of the 2004 Winter Simulation Conference. IEEE; 2004. A large deviations perspective on ordinal optimization; pp. 577–585. [Google Scholar]
- 98.Pasupathy R., Hunter S.R., Pujowidianto N.A., et al. Stochastically constrained ranking and selection via SCORE. ACM Trans. Model. Comput.Simul. (TOMACS) 2014;25(1):1–26. [Google Scholar]
- 99.Peng Y., Chen C.-H., Fu M.C., et al. Dynamic sampling allocation and design selection. INFORMS J. Comput. 2016;28(2):195–208. [Google Scholar]
- 100.Peng Y., Chong E.K.P., Chen C.-H., et al. Ranking and selection as stochastic control. IEEE Trans. Autom. Control. 2018;63(8):2359–2373. [Google Scholar]
- 101.Chick S.E., Inoue K. New two-stage and sequential procedures for selecting the best simulated system. Oper. Res. 2001;49(5):732–743. [Google Scholar]
- 102.Frazier P.I., Powell W.B., Dayanik S. A knowledge-gradient policy for sequential information collection. SIAM J. Control Optim. 2008;47(5):2410–2439. [Google Scholar]
- 103.Ryzhov I.O. On the convergence rates of expected improvement methods. Oper. Res. 2016;64(6):1515–1528. [Google Scholar]
- 104.Chen C.-H., He D., Fu M.C., et al. Efficient simulation budget allocation for selecting an optimal subset. INFORMS J. Comput. 2008;20(4):579–595. [Google Scholar]
- 105.Zhang G., Chen B., Jia Q.-S., et al. Efficient sampling policy for selecting a subset with the best. IEEE Trans. Autom. Control. 2023;68(8):4904–4911. [Google Scholar]
- 106.Zhang G., Peng Y., Zhang J., et al. Asymptotically optimal sampling policy for selecting top-m alternatives. INFORMS J. Comput. 2023;35(6):1261–1285. [Google Scholar]
- 107.Gao S., Chen W. Efficient feasibility determination with multiple performance measure constraints. IEEE Trans. Autom. Control. 2016;62(1):113–122. [Google Scholar]
- 108.Shi Z., Peng Y., Shi L., et al. Dynamic sampling allocation under finite simulation budget for feasibility determination. INFORMS J. Comput. 2022;34(1):557–568. [Google Scholar]
- 109.Peng Y., Chen C.-H., Fu M.C., et al. Efficient sampling allocation procedures for optimal quantile selection. INFORMS J. Comput. 2021;33(1):230–245. [Google Scholar]
- 110.Shin D., Broadie M., Zeevi A. Practical nonparametric sampling strategies for quantile-based ordinal optimization. INFORMS J. Comput. 2022;34(2):752–768. [Google Scholar]
- 111.Goodwin T., Xu J., Celik N., et al. Real-time digital twin-based optimization with predictive simulation learning. J. Simul. 2022:1–18. [Google Scholar]
- 112.Li H., Lam H., Peng Y. Efficient learning for clustering and optimizing context-dependent designs. Oper. Res. 2024;72(2):617–638. [Google Scholar]
- 113.Shen H., Hong L.J., Zhang X. Ranking and selection with covariates for personalized decision making. INFORMS J. Comput. 2021;33(4):1500–1519. [Google Scholar]
- 114.Hunter S.R., Nelson B.L. Advances in Modeling and Simulation: Seminal Research from 50 Years of Winter Simulation Conferences. Springer; 2017. Parallel ranking and selection; pp. 249–275. [Google Scholar]
- 115.Pei L., Nelson B., Hunter S. Parallel adaptive survivor selection. Oper. Res. 2024;72(1):336–354. [Google Scholar]
- 116.Zhong Y., Hong L.J. Knockout-tournament procedures for large-scale ranking and selection in parallel computing environments. Oper. Res. 2022;70(1):432–453. [Google Scholar]
- 117.Z. Zhang, Y. Peng, Sample-efficient clustering and conquer procedures for parallel large-scale ranking and selection, arXiv: 2402.02196 (2024).
- 118.R. Zhou, L.J. Hong, Y. Peng, AlphaRank: an artificial intelligence approach for ranking and selection problems, arXiv: 2402.00907 (2024).
- 119.Zhang G., Peng Y., Xu Y. Proceedings of the 2022 Winter Simulation Conference. IEEE; 2022. An efficient dynamic sampling policy for Monte Carlo tree search. [Google Scholar]
- 120.Liu X., Lam H., Peng Y. Training deep Q-network via Monte Carlo tree search for adaptive bitrate control in video delivery. Available at SSRN 4297671. 2022 [Google Scholar]
- 121.Preil D., Krapp M. Artificial intelligence-based inventory management: a Monte Carlo tree search approach. Ann. Oper. Res. 2022;308:415–439. [Google Scholar]
- 122.T. Ren, R. Zhou, J. Jiang, et al. RiskMiner: Discovering formulaic Alphas via risk seeking Monte Carlo tree search, preprint at arXiv: 305.08960 (2024).
- 123.Prashanth L.A., Fu M.C. Risk-sensitive reinforcement learning via policy gradient search. Found. Trends Mach. Learn. 2022;15(5):536–690. [Google Scholar]








