Abstract
Complex computer codes are frequently used in engineering to generate outputs based on inputs, which can make it difficult for designers to understand the relationship between inputs and outputs and to determine the best input values. One solution to this issue is to use design of experiments (DOE) in combination with surrogate models. However, there is a lack of guidance on how to select the appropriate model for a given data set. This study compares two surrogate modelling techniques, polynomial regression (PR) and kriging-based models, and analyses critical issues in design optimisation, such as DOE selection, design sensitivity, and model adequacy. The study concludes that PR is more efficient for model generation, while kriging-based models are better for assessing max-min search results due to their ability to predict a broader range of objective values. The number and location of design points can affect the performance of the model, and the error of kriging-based models is lower than that of PR. Furthermore, design sensitivity information is important for improving surrogate model efficiency, and PR is better suited to determining the design variable with the greatest impact on response. The findings of this study will be valuable to engineering simulation practitioners and researchers by providing insight into the selection of appropriate surrogate models. All in all, the study demonstrates surrogate modelling techniques can be used to solve complex engineering problems effectively.
Keywords: Design of experiment (DOE), Polynomial regression (PR), Kriging-based model, -Surrogate model, Underground shelter
1. Introduction
Engineering optimisation deals with the search for an optimal design of a system or element. The use of design optimisation techniques in engineering has been growing rapidly. In spite of the steady advance in computing architecture, the computational cost needed for carrying out a great number of simulations (due to many design variables) is still a major concern amongst the design engineers. Moreover, this approach often addresses one design at a time, whereby a designer might not be able to make a good design decision as they might fail to discover the functional relationship between the design variables and the responses [1]. In order to circumvent this problem, approximation model or surrogate model [2] can be adopted. There are numerous surrogate models available nowadays. Amongst these models, the Polynomial Regression (PR) [[3], [4], [5]] and the Kriging-based models [[6], [7], [8], [9]] are the popular ones in engineering [10,11]. More information regarding the application of surrogate models in engineering optimisation can be found in Refs. [1,10,12].
Recently, the number of studies focusing on applying different surrogate models (e.g., PR and Kriging-based) in optimisation has increased. However, it seems that the choice of surrogate model is very problem-dependent, and even for the same problem, small variations in the DOE can change the ranking of the methods [13,14]. Some comparative studies on PR and Kriging-based models in design optimisation have been reported [15,16]. Their modelling accuracies have been compared as well [17]. On the other hand [18], studied the advantages and disadvantages of surrogate models such as PR, Kriging, Multivariate Adaptive Regression Splines (MARS), and Radial Basis Function (RBF). The design sensitivities of PR and Kriging-based models have been reported, whereby these surrogate models were enhanced and the best strategy for model training was presented [19]. Also, the error measurements for PR and Kriging-based models in terms of noise-free functions have been investigated [11]. Some detailed reviews of the literature mentioned above are tabulated in Table 1 to provide some useful insights of design optimisation using various surrogate models. Despite the availability of numerous surrogate models, the choice of the appropriate model for a given dataset is still lacking, which raises concerns among design engineers.
Table 1.
An overview of the comparative studies of different surrogate models with indication of the test problem, experimental design, model choice, and design variables.
| Authors (year) | Test Problem | Experimental Design | Model Choice | Design Variables |
|---|---|---|---|---|
| Goel et al. [11] | Cantilever beam, Branin-Hoo function, Camelback | LHDs | PR, Kriging | 2-6 variables |
| Simpson et al. [15] | Aerospike nozzle problem | OA | PR, Kriging | 3 variables |
| Ahmed & Qin [16] | Hypersonic spiked blunt bodies | LHDs | PR, Exponential Kriging, Gaussian Kriging, General Exponential Kriging | 3 variables |
| Giunta & Watson [17] | High-Speed Civil Transport Aircraft | D-optimal | PR, Kriging | 5 and 10 variables |
| Jin et al. [18] | 13's Mathematical problem, Vehicle handling | LHDs | PR, MARS, RBF, Kriging | 2-16 variables |
| Rijpkema et al. [19] | Finite Elements Models | Full Factorial, LHDs | PR, Kriging | 3 variables |
| Devanathan & Koch [20] | Branin function, Michalewicz function, Ackley path function, | Full Factorial, OSFD, Random | PR, Kriging, Blind Kriging, RBF, EBF, Chebyshev | random |
LHDs = Latin Hypercube Sampling, PR = Polynomial Regression, OA = Orthogonal Array, MARS = Multivariate Adaptive Regression Splines, RBF = Radial Basis Function, OSFD = Optimal Space Filling Design, EBF = Elliptical Basic Function.
The choice of DOE, the sensitivity of the parameters to the response, and the suitability of the fitted surrogate model are a few crucial issues that must be addressed when using surrogate models for design optimisation. Barton [21] argues that addressing these concerns is necessary for ensuring the quality and dependability of the surrogate model. In this study, we aim to address the difficulty of understanding the relationship between inputs and outputs of complex computer codes by optimising a computational fluid dynamics (CFD) model of an underground shelter using PR and kriging-based models [22]. Their accuracy, sensitivity of design variables to response, model adequacy, and efficiency have been reported. In addition, the current results could guide the selection of an appropriate model for a given data set and demonstrate how surrogate modelling can be utilised to solve complex engineering problems.
The study focuses primarily on comparing the performance of PR and Kriging-based models as surrogate modelling approaches, rather than considering other popular methods such as Support Vector Regression (SVR) [23], Gaussian Process Regression (GPR) [23], Radial Basis Function (RBF) [24] and Artificial Neural Networks (ANN) [25,26]. The decision to exclude these alternative methods can be attributed to several factors. First, the choice of modelling techniques in a research study is usually determined by the context in which the study is conducted. PR and Kriging-based models [4,5,8,27,28] were chosen because of their suitability to the problem under study and their widespread use in the field. Secondly, by focusing exclusively on PR and Kriging-based models, the study is able to conduct a more detailed and focused comparative analysis and explore in depth the strengths, weaknesses and design optimisation performance of these two methods. Finally, it is important to note that the exclusion of other methods does not diminish their value or relevance. Different surrogate modelling techniques have their own advantages and may be more suitable for different scenarios.
Overall, it is obvious that further comprehensive studies of surrogate models are needed to establish clear selection criteria for different datasets. Such studies can help overcome current limitations and ensure that surrogate models are widely used in design optimisation.
2. Related work
Surrogate modelling techniques [[29], [30], [31], [32], [33]] have gained popularity in optimising engineering design problems, with researchers focusing on different applications. Koziel and Pietrenko-Dabrowska [34,35] contribute to this field by addressing the challenges of designing and optimising high-frequency systems, especially antennas and microwave circuits. They emphasise the importance of incorporating automation and optimisation techniques into the design to achieve optimal performance while addressing conflicting objectives such as electrical properties, physical size and cost. To overcome the computational cost of full-wave electromagnetic analysis, the authors propose the use of surrogate modelling techniques.
Koziel and Pietrenko-Dabrowska [36]also explore several methods to improve the effectiveness of surrogate modelling, including triangulation of reference designs, nested kriging and explicit dimensionality reduction. These techniques contribute to accurate and efficient surrogate models by encompassing the Pareto front with small regions of parameter space. The authors also focus on the application of surrogate modelling techniques to inverse problems in antenna design. They propose a framework for global optimisation of multiband antennas by combining metaheuristic techniques, surrogate modelling and sequential sampling methods. Their approach includes a knowledge-based inverse surrogate that accounts for essential response characteristics and enables cost-effective design optimisation.
Furthermore, Koziel and Pietrenko-Dabrowska [37] emphasise in their other work, the consideration of manufacturing tolerances and uncertainties in antenna design. They present an approach that uses surrogate modelling techniques for design centering of multiband antennas. By establishing a functional relationship between response characteristics and antenna geometry parameters through knowledge-based inverse regression models, the authors enable the prediction of parameter vectors that increase the probability of meeting design requirements under uncertainty. This approach significantly reduces the computational costs associated with the design centering process.
In addition to using surrogate models, researchers have explored the combination of machine learning (ML) [23,[38], [39], [40], [41], [42], [43]] genetic algorithms (GA) [23,38,[43], [44], [45], [46], [47]] design of experiments (DoE) [13,48] and computational fluid dynamics (CFD) [4,5,8,12,27,28,49] in engineering optimisation. Moiz et al. [43] presented an integrated ML and GA approach to optimising internal combustion engines, achieving comparable results to traditional CFD-GA approaches with significant time and cost savings. Badra et al. (2021) [39] developed a machine learning approach (ML-GGA) for optimising compression ignition petrol engines, which has higher accuracy and robustness compared to ML-GA. Badra et al. [40] conducted a comprehensive study to optimise the combustion system of a compression ignition engine running on commercial petrol and achieved significant improvements in engine performance and emissions with their approach ML-GGA. Owoyele et al. [50] proposed an automated surrogate-based optimisation approach called AutoML- GA for internal combustion engines that addresses the challenges of selecting optimal ML hyperparameters and unknown training data requirements.
In summary, the studies discussed highlight the benefits of using surrogate modelling, ML-GA and Auto ML-GA approaches to design optimisation. By incorporating advanced computational techniques such as surrogate modelling, ML, GA and CFD, these methods provide efficient and cost-effective solutions for optimising designs considering multiple objectives and uncertainties. They provide valuable insights and strategies for improving performance and reducing computational effort in engineering optimisation. However, this paper focuses specifically on the comparison of PR and kriging-based surrogate models in design optimisation, focusing on their strengths and weaknesses and the integration of CFD simulations.
3. Surrogate model and optimisation structure
Since the past few decades, the surrogate model approach has been adopted by many Heating Ventilation and Air-Conditioning (HVAC) practitioners. Each surrogate model is associated with numerical fitting. For instance, the PR method uses the Least Squares Curve Fit (LSCF) method while the Kriging-based method uses the Best Linear Unbiased Predictor (BLUP) method. The construction of PR or Kriging-based model depends on the location of design points (in a design space). This design space can be defined by the lower and upper bounds of the independent parameters, thereby forming a -dimensional cube. More explanation on these surrogate models can be found in Ref. [1].
3.1. Polynomial regression (PR)
PR is known to be straightforward and computationally efficient. It can be easily constructed; however, it is less accurate compared to more advanced schemes such as the Kriging-based model. PR is classically adopted in physical experiments. Nowadays, PR is applied in numerical modelling as well. PR is a statistical modelling technique that uses regression analysis to create a polynomial approximation of the computer analysis code [4,5]. This regression formula can be expressed in any order; however, the first- and second-order forms are the most popular ones. A second-order polynomial (Eq. (1)) was employed in the current study:
| (1) |
Here, is the predicted response value, represents the unknown regression coefficients, and and refer to the design variables. Using the conventional least squares equation, the polynomial coefficient can be estimated by minimising the sum of squared deviations between the predicted response value and the actual response value . The equation can be expressed as follows (Eq. (2)).
| (2) |
In the situation where has linear independence, the existence of the is then guaranteed. The vector denotes the set of observed responses across all design points. Therefore, the following representation of the matrix (Eq. (3)) can be described accordingly.
| (3) |
In this context, which is defined as , stands as a representative modelling term within the quadratic polynomial. The evaluation of the predicted response values at a new point can be efficiently achieved by replacing the respective within the polynomial equation.
3.2. Kriging-based model
Kriging is also known as a non-parametric interpolation method for Design and Analysis of Computer Experiment (DACE) [6,7]. In contrast to PR, Kriging was initially adopted for computer experiments via deterministic errors. The basic equation for the Kriging-based model (Eq. (4)) can be written as:
| (4) |
Here, is a low-order polynomial that interpolates the design points. However, as reported by Sacks [7], the term in Eq. (4) is typically taken as a constant value for modelling complex input-output relations in which the output can be viewed as . The equation can be expressed as follows (Eq. (5)).:
| (5) |
Here, is an unknown constant to be estimated from the observed response value, while is a realization of the normal distributed Gaussian stochastic function assumed to have zero mean, spatial covariance of the process variance and the correlation matrix function . The covariance matrix of (Eq. (6)) can be written as:
| (6) |
Here, is the symmetric correlation matrix with values of unity along the diagonal. Then, the correlation function is given as an exponential correlation function [6] written as (Eq. (7)):
| (7) |
Here, and refer to two design points, refers to the design parameter, refers to the number of design parameters, and refers to the unknown correlation parameters, which can be estimated from the Maximum Likelihood Estimation (MLE) as described by Booker [6]. Then, the unknown constant and the process variance can be calculated using the generalized least squares. The equation can be expressed as Eqs. (8), (9):
| (8) |
and
| (9) |
Here, is an matrix of design points depending on the choice of . However, the results of and depend on the unknown correlation parameters . Once the design parameters and correlation function are estimated from the best linear unbiased prediction method, the values of the predicted response at a new point can be computed from Eq. (10):
| (10) |
where is the correlation vector between and all design points.
3.3. Screening optimisation approach
The optimisation sequence is the ultimate aim of performing a surrogate model analysis in engineering problems. The two surrogate modelling techniques used in the current study involved analyses of several responses. As such, the Shifted Hammersley Sampling [51] was used for optimising the surrogate model responses because this sampling method is computationally efficient. The conventional Hammersley Sampling algorithm is a quasi-random number generator. It has a very low discrepancy; therefore, it is suitable for quasi-Monte-Carlo simulations. As highlighted by Diwekar and Kalagnanam [51], the discrepancy is a quantitative measure for optimally approximating the sequence uniform distribution. In other words, due to the inherent properties of the Monte Carlo simulation, this sampling method provides an unbiased search space and better uniformity for the design space. Also, it does not have any effect on the dimensionality.
The Hammersley Sampling algorithm can be constructed using a radical inverse function. Any integer can be written in a sequence radix-R format:
| (11) |
Here, and the square brackets denote the integral part. Then, the radical inverse function is defined by revising the order of digits of Eq. (11) about the decimal point in Eq. (12):
| (12) |
Thus, the Hammersley points on the k-dimensional search space can be expressed as Eq. (13):
| (13) |
where are the first k-1 prime numbers and the Hammersley points are .
4. Design of experiment (DOE)
4.1. Central Composite Design (CCD)
Central Composite Design (CCD) is the most popular experimental design method for fitting a PR model. It is known as one of the classical experimental design methods. It has a five-level fractional factorial designs, which is developed for fitting a second-order polynomial. CCD involves a two-level factorial points () augmented by two-star points and center points [1] defined on each factor. In addition, this method provides an alternative to experimental design for full factorial design (3N) while dealing with PR models. Technically, the method requires an experiment number of , where is the number of variables and is the replicate number of the central point [1]. In contrast, CCD tends to consistently distribute the design points around the edge and center of the design space.
4.2. Box-Behnken Design (BBD)
Box-Behnken Design (BBD) is an independent three-level of second-order experimental design method in which it does not contain an embedded factorial design [5,52]. It is also one of the classic experimental designs originated from the theory of DOE, which is useful in building a second-order PR model for the response variable. The locations of the design points could be in the median of each design area. The BBD requires a smaller number of design points, increasing efficiency and cost-effectiveness compared to traditional experimental design methods. In contrast to CCD, the focus of BBD is primarily on the design points that lie at the boundaries of the design area. From a technical point of view, BBD requires several experiments denoted by , where is the number of variables and is the number of central point's [1].
4.3. Latin Hypercube Designs (LHDs)
Latin Hypercube Designs (LHDs) have been widely used in DACE, and it is also known as a stratified sampling technique. This method is an advanced form of the Monte Carlo Sampling which avoids clustering samples. In other words, the design does not replicate any row or column with other points. In addition, LHDs involve non-uniform distribution of points over the design space. This design is a complex statistical design, which can fit well with the Kriging-based model [53]. It has been found to be efficient in estimating the predicted response value [54]. Technically, it is targeted at the input space formed by the k-dimensional cube in n runs for the k factors ( More details on LHDs can be found in Refs. [[53], [54], [55]].
4.4. Optimal Space Filling Design (OSFD)
Optimal Space Filling Design (OSFD) is the extended form of LHDs for DACE. It is also known as the Optimal Latin Hypercube [54]. In general, the LHD approach is optimised through several iterations. A more uniform space distribution of design points can be achieved via maximizing the distance between points (without point-sharing in rows or columns) [54,56]. This method is effective for complex surrogate model techniques such as Kriging-based model and Neural Network [48]. It seems that the coverage of OSFD is better than that of LHDs when dealing with a more substantial number of design points. Hence, it is good in modelling the actual behaviour of response influenced by many design variables.
5. Case study description
The current case study deals with the optimal design of ventilation shafts for ventilating an underground shelter [49]. The ventilation shafts consist of inlet shaft, outlet shaft and elbow shaft. Basically, the ventilation shaft is widely used in channelling fresh air into the shelter. As argued by Shetabivash [57] and Etheridge [58], the size of the opening has a significant influence on the airflow pattern inside the building. Edward and Randall [59] reported that inlet and outlet ventilation shafts of equal opening size would produce the highest ventilation rate per unit area. A similar argument was also presented by Andersen [60], in which the author found that the ventilation rate was dependent on the inlet or outlet areas as well as the vertical opening distance due to the buoyancy effect. Thus, the relationship between the ventilation rate and the opening areas must be considered when designing a ventilation system for an underground shelter.
5.1. CFD configuration, model geometry and mesh generation
In this study, the commercial software Computational Fluid Dynamics (CFD) based on the finite volume principle (namely ANSYS Fluent) was used to solve the Reynolds Averaged Navier-Stokes (RANS) equations. These equations describe the turbulent flow dynamics in the shelter. The experimental results [22] were used to validate the current CFD model. This shelter is a single-occupancy emergency shelter. A three-dimensional computer-aided design (CAD) of the experimental model was created specifically for the simulation (see Fig. 1). As depicted in Fig. 1b, the upper rectangular area was used to replicate the atmospheric circumstances of the actual experiment. The use of this condition allows the induction of air circulation through the inlet duct entrance, mimicking the authentic experimental environment. The heating cable was designed to replicate the heat release of a person in an enclosed space. Detailed description of the test case can be found in Ref. [28]. An unstructured tetrahedral mesh was then used to discretize the model (see Fig. 2a). Finer meshes were used in critical areas such as the ventilation shaft and the heating cable (see Fig. 2b). Six models with different element sizes were developed and simulated to investigate the effects of mesh size on flow results. The results are discussed in 4.3.
Fig. 1.
Schematic drawing of the underground shelter in this study.
Fig. 2.
Computational surface grid of the simulation model.
5.2. Computational technique and boundary conditions
The flow was considered stationary and incompressible. The realizable model was used to represent the flow turbulence as it can accurately reproduce the characteristics of the flow in an indoor setting [58]. The pressure-velocity coupling algorithm, i.e. SIMPLEC, was used. Furthermore, the use of higher order convective schemes is required to ensure accurate flow [8,28]. To ensure both accuracy and stability, all governing equations were discretised with second-order upwind schemes. The governing equations such as mass (14), momentum (15) and energy (16) conservation equations are:
| (14) |
| (15) |
| (16) |
In this context, signifies density, denotes time, is the velocity vector, corresponds to pressure, is the viscous stress tensor, refers to gravitational acceleration, encapsulates total energy, is indicative of the thermal conductivity of the fluid, is the absolute temperature, represents the source term and refers to viscous dissipation. The term on the right-hand side (RHS) of Eq. (15) symbolises the buoyancy force, where is the thermal expansion coefficient, is the reference density of the flow and is the operating temperature. The integration of the buoyancy model for the simulation of indoor air flows is documented in Ref. [61].
The simulation was then run for about 850 iterations to obtain a solution that showed convergence. The solution was considered to have converged when there were no longer any significant deviations in variables such as velocity, energy and turbulence (the scaled residual errors RMS dropped to 10−4) and the area had a net imbalance of less than 1%. To speed up the convergence of the solution, the Algebraic Multi-Grid (AMG) technique was used. The aggregation of the mass flow rate at the outlet of the opening shaft allowed the calculation of the post-simulation ventilation rate.
The boundary conditions were configured to match the experimental conditions as described in Ref. [22]. An invariant wind profile was specified at the domain inlet, while a zero static pressure was determined at the domain outlet. A free-slip condition was used for the domain wall (under atmospheric conditions). The surfaces of the ventilation shafts were treated as adiabatic, no-slip walls. The heat flux emanating from the heating cable was set to a value of 70 W/m2 [62,63]. The computational domain is shown in Fig. 3 and the boundary conditions are summarised in Table 2.
Fig. 3.
Computational domain of boundary conditions.
Table 2.
Boundary conditions for the simulation.
| Domain Inlet | constant velocity, 2.68 m/s at a constant temperature, |
|---|---|
| Domain Outlet | atmospheric pressure |
| Domain Wall | free slip wall at a constant temperature, |
| Inlet Shaft | adiabatic, no-slip walls () |
| Outlet Shaft | adiabatic, no-slip walls () |
| Heating Cable | heat flux, 70 W/m2 |
| Underground Shelter | adiabatic, no-slip walls () |
5.3. Grid independent test and CFD validation
Before the optimisation process of the Computational Fluid Dynamics (CFD) model, it is essential to compare the results of CFD with previous experimental results [22]. The Grid Independence Test (GIT) was first performed with six different grids (refer to Table 3). The range of ventilation rate deviation was between 3.35% and 14.05%, as shown in Table 3. The results generated with grids E and F do not seem to show any significant deviation. A corresponding result can be seen in Fig. 4, where the cases with larger mesh sizes (grids E and F) show analogous values for the axial velocity at the outlet of the opening shaft. Thus, when the number of cells reaches 1.63 million, the solution can be considered grid independent. The model with grid E was then selected for further flow analysis.
Table 3.
Grid parameters for six different mesh sizes.
| Grid | A | B | C | D | E | F |
|---|---|---|---|---|---|---|
| Element Size | Default | 0.5 | 0.2 | 0.175 | 0.15 | 0.125 |
| Number of Cells | 759,245 | 772,180 | 1,113,780 | 1,291,880 | 1,625,121 | 2,390,589 |
| 11.95% | 14.05% | 13.00% | 6.5% | 3.56% | 3.35% |
Fig. 4.
Verification charts.
A comprehensive validation of the model CFD shown in Fig. 1 was carried out. The calculated ventilation rate was 0.0456 m3/s, in contrast to the measured ventilation rate of 0.0477 m3/s [22]. The percentage discrepancy between these two values was only 4.4%. Consequently, the current CFD model can be trusted to accurately predict the ventilation efficiency of naturally ventilated underground shelters. Detailed validation studies for the current model CFD can be found in Ref. [28].
6. Surrogate model methodology
The model can be optimised once the CFD model is validated. Here, the ANSYS Design Exploration 16.0 software was used to solve the selected DOE and surrogate models [13]. Fig. 5 shows the optimisation process in this study. The first step of parametric and optimisation analysis consists of identifying the objective functions, design variables (input parameters), and constraints. In this study, the objective function is to maximize the ventilation rate. The selection of design variables is influenced by the findings of previous studies (presented in Section 4). The inlet opening (P1), the ratio of radius to height of inlet shaft (P2), and the outlet opening (P3) were expected to have a significant influence on the ventilation rate (response) (P4) required for one person inside the underground shelter (see Fig. 6). The ranges of the design variables (i.e. lower bound, upper bound and constraint) were reported in Ref. [22]. The design variables including their ranges are summarised in Table 4.
Fig. 5.
Flowchart of the optimisation process.
Fig. 6.
Design variables.
Table 4.
Types of design variables with upper and lower boundary parameters.
| Parameters | Name | Upper Bound (m) | Lower Bound (m) | Constraints | |
|---|---|---|---|---|---|
| Factors (IV) | P1 | Inlet Opening | 0.1524 | 0.1016 | – |
| P2 | Ratio of Elbow Shaft | 2.50 | 0.75 | 0.75 | |
| P3 | Outlet Opening | 0.1524 | 0.1016 | – | |
| Response (DV) | P4 | Ventilation Rate (m3/s) | |||
The selection of a proper DOE is significant [13,14] while performing numerical experiments. In this study, the distribution of design points was determined by using four types of DOE, namely BBD, CCD, LHDs and OSFD. These types of DOE have been recommended for PR and Kriging-based models [3,48,53]. In this study, a DOE approach was used to efficiently locate the data points and improve the computational time and accuracy of the surrogate model. For the DOE analyses conducted in this study, a total of 14–15 data points were generated, each representing varying input parameters for the Computational Fluid Dynamics (CFD) model. In the current work, the accuracy of response prediction was expressed using the coefficient of determination R2 (i.e. Adjusted R2 and Predicted R2). In addition, the early prediction of the maximum and minimum search of ventilation rate was also investigated.
In this study, both PR and Kriging-based models were adopted for the purposes of regression and interpolation analyses, respectively. In this stage, the relationship between the input parameters and the response was investigated by examining the local sensitivity curves [64]. Local or global sensitivity analyses depend on the objectives of the study, the resources and the complexity of the model. Both methods have their advantages, but local sensitivity analysis is more appropriate for this study. Complex models make global sensitivity analysis computationally difficult and time-consuming [14,65] Local sensitivity analysis, on the other hand, focuses on a specific region of parameter space or a collection of variables, making it computationally more efficient [39,66]. This localised technique is useful for answering specific hypotheses that require a deep understanding of the model's behaviour in a particular domain. It allows targeted studies of how minor changes in input parameters affect model results in that region [8]. Local sensitivity analysis also helps prioritise factors by validating or quantifying their influence. When computational resources or time are limited, this strategy reduces the scope of the analysis and saves work. Also, in this study contour plot was used to visualize the response of the ventilation rate due to the input parameters. The plot outlines the minimum and maximum range of the response data. In fact, it also gives an early prediction of the possibility of the highest ventilation rate before the optimisation result is finalised. Table 5 presents the surrogate model techniques employed in this study.
Table 5.
Surrogate model techniques.
| Experimental Design | Model Choice | Model Fitting | Sample Approximation Technique |
|---|---|---|---|
| CCD | Second Order Polynomial | Least square regression | PR |
| BBD | |||
| LHDs | Realization of a stochastic process | Best linear unbiased predictor (BLUP) | Kriging |
| OSFD |
The final step involves the selection of objective function and screening optimisation. As mentioned earlier, the objective function involves the maximization of ventilation rate. The sample set can be generated by screening in order to determine the effect of input parameters on the output parameters. Technically, there are three optimisation methods available in ANSYS Design Explorer, i.e. Shifted Hammersly Sampling (SHS), Multi-objective Genetic Algorithm (MOGA) and Non-linear Programming by Quadratic Lagrangian (NLPQL). SHS is a non-iterative direct sampling method using a quasi-random number generator based on the Hammersley algorithm which is good for obtaining the preliminary optimal solution. MOGA is a more advanced approach which is usually employed when there are multiple objective functions. NLPQL is a fast gradient-based local optimisation algorithm for single objective function. In the current study, we have decided to use SHS method which is a simple approach based on sampling and sorting. In fact, it supports multiple objective functions and constraints as well as all types of input parameter [67]. Moreover, this approach is able to provide a global overview of the design space and to identify the local minima.
7. Result and discussion
7.1. Sensitivity analysis
A correlation matrix was created to observe the response and its correlation with other parameters (factors). In other words, the strong and weak correlations of each factor with the design response should be determined. Pearson's rank correlation was used for this purpose, which requires 100 samples with 5% of the corresponding threshold for filtering along the correlation value [8]. The correlation matrix, as shown in Fig. 7, provides a visual representation of the correlations, allowing the user to see which factors have the strongest relationship with the response variable. This information can be used to prioritise which factors to focus on for optimisation or further investigation. The colour coded matrix that resulted indicated the strength of the correlation between each factor and the response variable. When the value is closer to the absolute value of 1, a stronger relationship is expected. According to the matrix, P3 has the most influence on P4 with a correlation value of 0.7704, followed by P1 with a correlation value of 0.6078 and P2 with a correlation value of only 0.0223. A scatter plot in Fig. 8 was also created to produce both linear and quadratic trend lines for the most highly correlated factors. The quadratic trend line performed better for each factor in this case, indicating a non-linear relationship between the factors and the response variable. As shown in Fig. 8 the highest estimated coefficient of determination (R2), with a percentage of 61.25%, is found for the correlation between P3 and P4 (see Fig. 8c). The coefficient of determination (R2) for the correlation between P1 and P4 is only 37.16% (see Fig. 8a), while the correlation between P2 and P4 has a lower coefficient of determination (R2) of 0.15% (see Fig. 8b). Overall, P1 and P3 showed a stronger correlation with a higher coefficient of determination (R2) than P4. The correlation matrix and scatter plot information can be used to prioritise which factors to focus on for optimisation or further investigation.
Fig. 7.
Correlation matrix.
Fig. 8.
Correlation charts.
7.2. Comparison between the PR and kriging-based models
The simulation results of the PR and Kriging-based models can be compared by plotting the predicted response values against the observed values. The final surrogate model using different DOE methods is shown in Table 6. Typically, a good DOE model for each surrogate model would exhibit high Adjusted R2 and Predicted R2 values. Technically, the adjusted R2 value is the modification of R2 (the ratio of the model sum of squares to the total sum of squares) for the number of description terms in the model while the predicted R2 is used to identify the prediction of the model response based on the new observations and to give the first insight into the sensitivity of the design solution. Also, in the current study, we have identified the minimum and maximum search outputs of the response to yield a better prediction for the improvement of optimisation.
Table 6.
PR and Kriging surrogate model with the best quality of fit using different DOE methods.
| Surrogate model | DOE | Design Point | Max-min Search |
Adj R2 | Pred R2 | Quality of Fit | |
|---|---|---|---|---|---|---|---|
| Min Output (m3/s) | Max Output (m3/s) | ||||||
| PR |
BBD | 14 | 0.0192 | 0.0468 | 0.996 | 0.998 | ![]() |
| CCD |
15 |
0.0189 |
0.0466 |
0.978 |
0.983 |
![]() |
|
| Surrogate model |
DOE |
Design Point |
Max-min Search | RMSE |
Quality of Fit |
||
| Min Output (m3/s) |
Max Output (m3/s) |
||||||
| Kriging-based | LHDs | 15 | 0.0179 | 0.0476 | 1.3 × 10−9 | ![]() |
|
| OSFD | 15 | 0.0167 | 0.0471 | 1.5 × 10−5 | ![]() |
||
As shown in Table 6, he PR method with BBD generates a smaller number of design points, which may result in less accurate model predictions due to a lower sampling density of the design space. PR may thus be appropriate for simpler systems with lower computational costs. Furthermore, PR assumes that the data is free of noise, which may result in inaccurate predictions in the presence of data uncertainties. When evaluating the max-min search result, the Kriging approach produces 0.0167 m3/s and 0.0476 m3/s, respectively. The points developed in the design space show that the Kriging-based approach is superior to PR. This is because kriging-based models are better at modelling nonlinear relationships and handling noisy data, which can lead to more accurate predictions in the presence of complex system behaviour and data uncertainties. Kriging-based approaches also perform well in well-sampled design spaces, where a higher number of design points is typically required to achieve higher surrogate model accuracy. In addition, kriging-based methods provide uncertainty estimates for the predicted values, which can help in evaluating the reliability of the model predictions. In general, the higher the number of design points, the higher the accuracy of the surrogate model. It appears that both LHDs and OSFDs produce greater coverage of points in the design space compared to CCD and BBD. It should be noted, however, that this result is only a preliminary prediction for determining the maximum and minimum likelihood of the result prior to performing the optimisation.
As noted from Table 6 (see Quality of Fit), the PR technique employs the Least Square Regression to fit the simulation results while the Kriging-based technique interpolates the simulation results and fits the Maximum Likelihood Estimation (MLE) as well as the Unbiased Linear Predictor-criterion (ULP) on the results. On the other hand, the polynomial model can be evaluated by computing the R2 error. In general, higher R2 indicates better Quality of Fit and accuracy. For instance, the PR + BBD model is more desirable, and it has a stronger correlation with the Quality of Fit (e.g. R2 = 0.998). Basically, the Quality of Fit is a statistical significance of the surrogate model coefficients, which describes how well the observations points are fitted. This method is known as local surrogate model-based optimisation. Meanwhile, for evaluating the interpolation model, cross-validation and Root Mean Square Error (RMSE) can be employed to give a more accurate approximation over the design points. This argument was also discussed in Refs. [15,68] and the method was known as global surrogate model-based optimisation. Since the Kriging-based techniques use cross-validation and RMSE, the integrity of the function should not be dependent on the values of Adjusted R2 and Predicted R2 because the more the parameters involved in the system, the higher the value of the regression sum of squares which might lead to value greater than 1.0 [69]. Indeed, in current study, the values of Adjusted R2 and Predicted R2 were equal to 1.0, as ANSYS automatically truncates those values that are above 1.0 [67]. However, according to equation (17):
| (17) |
the current RMSE value varies from indicating that the regression function is adequate in representing the model. In other words, the closer the value of RMSE is to zero, the better the quality of fit. In fact, the current error estimations were obtained based on bias error rather than noise (as the error is stemming from model inadequacy instead of noise). Thus, it seems that Kriging-based model is more accurate than PR model. More explanations on the comparison between PR and Kriging-based models in terms of error measurement of noisy response can be found in Ref. [11].
For the sensitivity and response point of the design solution, it seems that PR has better sensitivity factors and higher response points as compared to the Kriging-based approach (see response point). Table 7 shows the variation of output with respect to the input parameters. From the results, it seems that the effects of P1 and P3 on the ventilation rate are more apparent than that of the ratio of elbow shaft (P2). Higher P1 or P3 would increase the ventilation rate. On the other hand, an opposite trend is observed for P2. Essentially, the local sensitivity study was conducted to examine the weight of each parameter around the response point. Each curve represents the impact of an enabled input parameter towards the output. Thus, PR is the best method to estimate interaction and even quadratic effects of the shape of response surface as PR was initially developed to analyse the cause-effect relationship of a physical experiment [3]. On top of that, it seems that both PR and Kriging-based approach reflect an important contribution of P1 and P3 on P4 which both of it detect a stronger effect from P3 (higher slope of the curves for P3). These findings are in qualitative agreement with the correlation matrix and charts (see Fig. 7, Fig. 8).
Table 7.
PR and Kriging surrogate model with the best sensitivity and response point using different DOE methods.
| Surrogate model | DOE | Inlet | Outlet | Response Point | Local Sensitivity Curves |
|---|---|---|---|---|---|
| PR |
BBD | 0.326 | 0.537 | (0.5, 0.0309) | ![]() |
| CCD |
0.397 |
0.509 |
(0.5, 0.0307) |
![]() |
|
| DOE |
Inlet |
Outlet |
Response Point |
Local Sensitivity Curves |
|
| Kriging-based | LHDs | 0.353 | 0.533 | (0.5, 0.0300) | ![]() |
| OSFD | 0.359 | 0.519 | (0.5, 0.0301) | ![]() |
The Kriging-based approach, however, was developed to fix the random error produced by a computer experiment [7]. In addition to the sensitivity factors and response point, a graphical comparison between PR and Kriging-based approaches was conducted to visualize the highest ventilation rate predicted by all types of DOE. Fig. 9, presents the contour plots of P1, P3, and P4. All results are almost similar, where the highest ventilation rate can be identified at the top right corner of the contour plot between the P1 and P3 factors. As seen in Fig. 9, the Z-axis represents the predicted ventilation rates (P4) of PR and Kriging-based models, while the X- and Y- axes represent the strong correlation of factors (e.g. P1 is Inlet Opening, and P3 is Outlet Opening, respectively). Overall, the design point distribution depends on type of DOE. The DOEs used in the PR approach (e.g. CCD (see Fig. 9a) and BBD (see Fig. 9b)) tend to generate the design points near the boundary of a design space. On the other hand, the DOEs used in the Kriging-based approach (see Fig. 9c and d) are known to yield a more uniform distribution of design points over the design space [70]. Therefore, they are good in predicting a larger dispersion of objective values. This argument has been verified, as presented in Table 6, for which the Kriging-based approach can identify the highest and the lowest max-min search outputs of the ventilation rate due to the dispersed locations of the design points.
Fig. 9.
Contour plot of different DOE.
7.3. Screening optimisation results
The screening approach (Shifted Hammersley Sampling) was employed for each surrogate model. It is one of the optimisation methods used in the ANSYS Design Explorer. A total of 1000 distributed sample sets were generated to find the optimal design based on the objective function. The results of the screening method are summarised in Table 8. It can be seen that the predicted maximum ventilation rates using PR and Kriging-based models increase by 6.14% and 4.39% (relative to the initial design), respectively. To validate the optimised result, the predicted optimal design is modelled using CFD. Table 9 compares the values obtained from the Shifted Hammersley Sampling and CFD analysis. By examining the results, the maximum error is within 2%–4%, which is within the permissible range of the current problem.
Table 8.
Summary of optimisation results.
| Variables |
Geometrical Dimensions |
Ventilation Rate (m3/s) |
|||
|---|---|---|---|---|---|
| P1 (m) |
P2 (m) |
P3 (m) |
|||
| Initial Design | 0.1524 | 1.67 | 0.1524 | 0.0456 | |
| PR | BBD | 0.1524 | 2.13 | 0.1524 | 0.0484 |
| CCD | 0.1524 | 2.13 | 0.1524 | 0.0484 | |
| Kriging-based | LHDs | 0.1524 | 2.31 | 0.1524 | 0.0476 |
| OSFD | 0.1524 | 2.31 | 0.1524 | 0.0470 | |
Table 9.
Summary of verification results.
| Variables | Predicted Optimisation (m3/s) | CFD (m3/s) | Error (%) | |
|---|---|---|---|---|
| PR | BBD | 0.0484 | 0.0465 | 4.08% |
| CCD | 0.0484 | 0.0465 | 4.08% | |
| Kriging-based | LHDs | 0.0476 | 0.0461 | 3.25% |
| OSFD | 0.0470 | 0.0461 | 2.39% | |
8. Conclusion
This paper compared two common surrogate modelling techniques, namely PR and Kriging-based models, for optimising a CFD model by identifying the functional relationship between the design variables and the response. PR models use a polynomial function to fit the simulation results, while Kriging-based models use an interpolation technique to fit the simulation results. Both techniques can be used for optimisation, but their effectiveness depends on the specific problem and the nature of the design space.
PR was discovered to be computationally efficient and appropriate for model building because it requires fewer design points, resulting in less computation time. The Kriging-based model, on the other hand, is a more suitable method for evaluating the max-min search performance of the response. This is probably due to the large number of design points, which is beneficial for predicting a larger dispersion of the objective values. The number and location of the design points may also affect the performance of the model created. When examining the quality of fit, PR appears to rely on estimating the R2 error to capture the correlation between the predicted and observed points. In this case, the Kriging-based model is superior because all predicted and observed points are fitted with a more accurate function than the design points. Furthermore, we discovered that the kriging-based model has a lower error than the PR model. In fact, the efficiency of a surrogate model can be improved by including information such as design sensitivity. In this case, PR models are better suited for estimating interaction and quadratic effects of the input parameters, while Kriging-based models are better suited for identifying the main effects of the input parameters. SHS was used to predict the preliminary optimal solution for optimisation, and the optimised result of PR is slightly higher than that of the kriging-based model. The current evaluation of surrogate model performance is thought to be useful for the design of HVAC components.
All in all, there is no one-size-fits-all approach to surrogate modelling for optimisation. Both PR and kriging-based models have advantages and disadvantages, and the choice between them is determined by the nature of the problem and the design space. PR is computationally efficient and appropriate for model building, whereas kriging-based models are better for evaluating the response's max-min search performance. Kriging-based models are also better at fitting predicted and observed points with a more accurate function than the design points. The selection of an optimisation procedure should be based on a careful assessment of the performance of each procedure to determine which method provides the best results for optimisation.
Author contribution statement
Azfarizal Mukhtar: Conceived and designed the experiments; Performed the experiments; Analyzed and interpreted the data; Contributed reagents, materials, analysis tools or data; Wrote the paper.
Ahmad Shah Hizam Md Yasir: Performed the experiments; Analyzed and interpreted the data.
Mohamad Fariz Mohamed Nasir: Contributed reagents, materials, analysis tools or data; Wrote the paper.
Data availability statement
Data will be made available on request.
Declaration of competing interest
The authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper.
References
- 1.Simpson T.W., Poplinski J.D., Koch P.N., Allen J.K. Metamodels for computer-based engineering design: survey and recommendations. Eng. Comput. 2001;17:129–150. doi: 10.1007/PL00007198. [DOI] [Google Scholar]
- 2.Kleijnen P.C. Sensitivity analysis of simulation experiments:regression analysis and statistical design. Math. Comput. Simulat. 1992;34:297–315. [Google Scholar]
- 3.Myers R.H., Montgomery D.C., Anderson-cook C.M. third ed. John Wiley & Sons Ltd; 2009. Response surface methodology: process and product optimization using designed experiments. [DOI] [Google Scholar]
- 4.Mukhtar A., Ng K.C., Yusoff M.Z. Passive thermal performance prediction and multi-objective optimization of naturally-ventilated underground shelter in Malaysia. Renew. Energy. 2018;123:342–352. doi: 10.1016/j.renene.2018.02.022. [DOI] [Google Scholar]
- 5.Mukhtar A., Yusoff M.Z., Khai Ching N., Mohamed Nasir M.F. Application of box-behnken design with response surface to optimize ventilation system in underground shelter. J. Adv. Res. Fluid Mech. Therm. Sci. 2018;52:161–173. [Google Scholar]
- 6.Booker A.J. 7th AIAA/USAF/NASA/ISSMO Symp. Multidiscip. Anal. Optim.; 1998. Design and analysis of computer experiments; pp. 118–128. [Google Scholar]
- 7.Sacks J., Welch W.J., Mitchell T.J., Wynn H.P. Design and analysis of computer experiments. Stat. Sci. 1989;4:409–435. [Google Scholar]
- 8.Mukhtar A., Ng K.C., Yusoff M.Z. Optimal design of opening ventilation shaft by kriging metamodel assisted multi-objective genetic algorithm. Int. J. Model. Optim. 2017;7:92–97. doi: 10.7763/IJMO.2017.V7.565. [DOI] [Google Scholar]
- 9.Koziel S., Pietrenko-Dabrowska A. Design-oriented modeling of antenna structures by means of two-level kriging with explicit dimensionality reduction. AEU - Int. J. Electron. Commun. 2020;127 doi: 10.1016/j.aeue.2020.153466. [DOI] [Google Scholar]
- 10.Wang G.G., Shan S. Review of metamodeling techniques in support of engineering design optimization. J. Mech. Des. 2007;129:370–380. doi: 10.1115/1.2429697. [DOI] [Google Scholar]
- 11.Goel T., Hafkta R.T., Shyy W. Comparing error estimation measures for polynomial and kriging approximation of noise-free functions. Struct. Multidiscip. Optim. 2009;38:429–442. doi: 10.1007/s00158-008-0290-z. [DOI] [Google Scholar]
- 12.Mukhtar A., Yusoff M.Z., Ng K.C. The potential influence of building optimization and passive design strategies on natural ventilation systems in underground buildings : the state of the art. Tunn. Undergr. Space Technol. 2019;92 [Google Scholar]
- 13.Pei Y., Zhang A., Pal P., Zhao L., Zhang Y., Som S. Artif. Intell. Data Driven Optim. Intern. Combust. Engines. Elsevier; 2022. Computational fluid dynamics–guided engine combustion system design optimization using design of experiments; pp. 103–123. [DOI] [Google Scholar]
- 14.Pei Y., Pal P., Zhang Y., Traver M., Cleary D., Futterer C., Brenner M., Probst D., Som S. SAE Tech. Pap. 2019-Janua. 2019. CFD-guided combustion system optimization of a gasoline range fuel in a heavy-duty compression ignition engine using automatic piston geometry generation and a supercomputer. [DOI] [Google Scholar]
- 15.Simpson T., Mistree F., Korte J., Mauery T. Comparison of response surface and kriging models for multidisciplinary design optimization, 7th AIAA/USAF/NASA/ISSMO Symp. Multidiscip. Anal. Optim. 1998:381–391. doi: 10.2514/6.1998-4755. [DOI] [Google Scholar]
- 16.Ahmed M.Y.M., Qin N. 13th Int. Conf. Aerosp. Sci. Aviat. Technol. 2009. Comparison of response surface and kriging surrogates in aerodynamic design optimization of hypersonic spiked blunt bodies; pp. 1–17. [Google Scholar]
- 17.Giunta A.A., Watson L.T. 7th AIAA/USAF/NASA/ISSMO Symp. Multidiscip. Anal. Optim. 1998. A comparison of approximation modeling techniques : polynomial versus interpolating models; pp. 1–13. [DOI] [Google Scholar]
- 18.Jin R., Chen W., Simpson T.W. Comparative studies of metamodelling techniques under multiple modelling criteria. Struct. Multidiscip. Optim. 2001;23:1–13. doi: 10.1007/s00158-001-0160-4. [DOI] [Google Scholar]
- 19.Rijpkema J., Etman L., Schoofs A. Use of design sensitivity information in response surface and kriging metamodels. Optim. Eng. 2001;2:469–484. [Google Scholar]
- 20.Devanathan S., Koch P.N. Comparison of meta-modeling approaches for optimization. Int. Mech. Eng. Congr. Expo. 2011:1–9. [Google Scholar]
- 21.Barton R.R. Proc. 2009 Winter Simul. Conf. 2009. Simulation optimization using metamodels; pp. 230–238. [Google Scholar]
- 22.King J.C. Port Hueneme; California: 1965. Gravity ventilation of underground shelters. [DOI] [Google Scholar]
- 23.Ghalandari M., Mukhtar A., Shah A., Yasir H. Thermal conductivity improvement in a green building with Nano insulations using machine learning methods. Energy Rep. 2023;9:4781–4788. doi: 10.1016/j.egyr.2023.03.123. [DOI] [Google Scholar]
- 24.Valadão M.A.C., Batista L.S. A comparative study on surrogate models for SAEAs. Opt Lett. 2020;14:2595–2614. doi: 10.1007/s11590-020-01575-2. [DOI] [Google Scholar]
- 25.Magnier L., Haghighat F. Multiobjective optimization of building design using TRNSYS simulations, genetic algorithm, and Artificial Neural Network. Build. Environ. 2009;45:739–746. doi: 10.1016/j.buildenv.2009.08.016. [DOI] [Google Scholar]
- 26.Stavrakakis G.M., Karadimou D.P., Zervas P.L., Sarimveis H., Markatos N.C. Selection of window sizes for optimizing occupational comfort and hygiene based on computational fluid dynamics and neural networks. Build. Environ. 2011;46:298–314. doi: 10.1016/j.buildenv.2010.07.021. [DOI] [Google Scholar]
- 27.Azman A., Yusoff M.Z., Mukhtar A., Gunnasegaran P., Ching N.K. Application of box- behnken design with response surface methodology to analyse friction characteristics for corrugated pipe via CFD. CFD Lett. 2023;7:1–13. [Google Scholar]
- 28.Mukhtar A., Ng K.C., Yusoff M.Z. Design optimization for ventilation shafts of naturally-ventilated underground shelters for improvement of ventilation rate and thermal comfort. Renew. Energy. 2018;115:183–198. doi: 10.1016/j.renene.2017.08.051. [DOI] [Google Scholar]
- 29.Koziel S., Pietrenko-Dabrowska A., Al-Hasan M. Low-cost multi-criteria design optimization of compact microwave passives using constrained surrogates and dimensionality reduction. Int. J. Numer. Model. Electron. Network. Dev. Field. 2021;34:1–12. doi: 10.1002/jnm.2855. [DOI] [Google Scholar]
- 30.Koziel S., Pietrenko-Dabrowska A., Mahrokh M. Globalized simulation-driven miniaturization of microwave circuits by means of dimensionality-reduced constrained surrogates. Sci. Rep. 2022;12:1–17. doi: 10.1038/s41598-022-20728-0. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 31.Koziel S., Pietrenko-Dabrowska A. Low-cost quasi-global optimization of expensive electromagnetic simulation models by inverse surrogates and response features. Sci. Rep. 2022;12:1–17. doi: 10.1038/s41598-022-24250-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 32.Calik N., Güneş F., Koziel S., Pietrenko-Dabrowska A., Belen M.A., Mahouti P. Deep-learning-based precise characterization of microwave transistors using fully-automated regression surrogates. Sci. Rep. 2023;13:1–16. doi: 10.1038/s41598-023-28639-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 33.Koziel S., Pietrenko-Dabrowska A. Tolerance optimization of antenna structures by means of response feature surrogates. IEEE Trans. Antenn. Propag. 2022;70:10988–10997. doi: 10.1109/TAP.2022.3187665. [DOI] [Google Scholar]
- 34.Koziel S., Pietrenko-Dabrowska A. Knowledge-based performance-driven modeling of antenna structures. Knowl. Base Syst. 2022;237 doi: 10.1016/j.knosys.2021.107698. [DOI] [Google Scholar]
- 35.Koziel S., Pietrenko-Dabrowska A. Global EM-driven optimization of multi-band antennas using knowledge-based inverse response-feature surrogates. Knowl. Base Syst. 2021;227 doi: 10.1016/j.knosys.2021.107189. [DOI] [Google Scholar]
- 36.Koziel S., Pietrenko-Dabrowska A. Recent advances in accelerated multi-objective design of high-frequency structures using knowledge-based constrained modeling approach. Knowl. Base Syst. 2021;214 doi: 10.1016/j.knosys.2020.106726. [DOI] [Google Scholar]
- 37.Koziel S., Pietrenko-Dabrowska A. Rapid design centering of multi-band antennas using knowledge-based inverse models and response features. Knowl. Base Syst. 2022;252 doi: 10.1016/j.knosys.2022.109360. [DOI] [Google Scholar]
- 38.Owoyele O., Pal P., Vidal Torreira A. An automated machine learning-genetic algorithm framework with active learning for design optimization. J. Energy Resour. Technol. 2021;143 doi: 10.1115/1.4050489. [DOI] [Google Scholar]
- 39.Badra J.A., Khaled F., Tang M., Pei Y., Kodavasal J., Pal P., Owoyele O., Fuetterer C., Mattia B., Aamir F. Engine combustion system optimization using computational fluid dynamics and machine learning: a methodological approach. J. Energy Resour. Technol. Trans. ASME. 2021;143:1–11. doi: 10.1115/1.4047978. [DOI] [Google Scholar]
- 40.Badra J., Khaled F., Sim J., Pei Y., Viollet Y., Pal P., Futterer C., Brenner M., Som S., Farooq A., Chang J. Combustion system optimization of a light-duty GCI engine using CFD and machine learning. 2020. [DOI]
- 41.Diego I., Torno S., Toraño J., Menéndez M., Gent M. A practical use of CFD for ventilation of underground works. Tunn. Undergr. Space Technol. 2011;26:189–200. doi: 10.1016/j.tust.2010.08.002. [DOI] [Google Scholar]
- 42.Eisenhower B., O'Neill Z., Narayanan S., Fonoberov V.A., Mezic I. A methodology for meta-model based optimization in building energy models. Energy Build. 2012;47:292–301. doi: 10.1016/j.enbuild.2011.12.001. [DOI] [Google Scholar]
- 43.Moiz A.A., Pal P., Probst D., Pei Y., Zhang Y., Som S., Kodavasal J. A machine learning-genetic algorithm (ML-GA) approach for rapid optimization using high-performance computing. SAE Int. J. Commer. Veh. 2018;11:291–306. doi: 10.4271/2018-01-0190. [DOI] [Google Scholar]
- 44.Wright J.A., Loosemore H.A., Farmani R. Optimization of building thermal design and control by multi-criterion genetic algorithm. Energy Build. 2002;34:959–972. doi: 10.1016/S0378-7788(02)00071-3. [DOI] [Google Scholar]
- 45.Li M., Li G., Azarm S. A kriging metamodel assisted multi-objective genetic algorithm for design optimization. J. Mech. Des. 2008;130 doi: 10.1115/1.2829879. [DOI] [Google Scholar]
- 46.Xue Y., Zhai Z.J., Chen Q. Inverse prediction and optimization of flow control conditions for confined spaces using a CFD-based genetic algorithm. Build. Environ. 2013;64:77–84. doi: 10.1016/j.buildenv.2013.02.017. [DOI] [Google Scholar]
- 47.Nassif N., Kajl S., Sabourin R. Optimization of HVAC control system strategy using two-objective genetic algorithm. HVAC R Res. 2005;11:459–486. doi: 10.1080/10789669.2005.10391148. [DOI] [Google Scholar]
- 48.Nasser M., Jawad B. ASME Int. Mech. Eng. Congr. Expo.; 2008. Design of Experiments as Effective Design Tools; pp. 1–8. [Google Scholar]
- 49.Mukhtar A., Ng K.C., Yusoff M.Z., Tey W.Y., Tan L.K. Performance assessment of passive heating and cooling techniques for underground shelter in equatorial climate. J. Adv. Res. Fluid Mech. Therm. Sci. 2019;61:20–32. [Google Scholar]
- 50.Owoyele O., Pal P., Vidal Torreira A., Probst D., Shaxted M., Wilde M., Senecal P.K. Application of an automated machine learning-genetic algorithm (AutoML-GA) coupled with computational fluid dynamics simulations for rapid engine design optimization. Int. J. Engine Res. 2022;23:1586–1601. doi: 10.1177/14680874211023466. [DOI] [Google Scholar]
- 51.Diwekar U.M., Kalagnanam J.R. Efficient sampling technique for optimization under uncertainty. AIChE J. 1997;43:440–447. doi: 10.1002/aic.690430217. [DOI] [Google Scholar]
- 52.Montgomery D.C. fifth ed. John Wiley & Sons, Inc.; 2002. Design and Analysis of Experiments. [DOI] [Google Scholar]
- 53.Kleijnen J.P.C. Kriging metamodeling in simulation: a review. Eur. J. Oper. Res. 2009;192:707–716. doi: 10.1016/j.ejor.2007.10.013. [DOI] [Google Scholar]
- 54.Park J.S. Optimal Latin-hypercube designs for computer experiments. J. Stat. Plann. Inference. 1994;39:95–111. doi: 10.1016/0378-3758(94)90115-5. [DOI] [Google Scholar]
- 55.Kleijnen J.P.C. Regression and kriging metamodels with their experimental designs in simulation: a review. Eur. J. Oper. Res. 2016;256:1–16. doi: 10.1016/j.ejor.2016.06.041. [DOI] [Google Scholar]
- 56.Santer T.J., Williams B.J., Notz W.I. Springer Science & Business Media; New York: 2003. The Design and Analysis of Computer Experiments. [Google Scholar]
- 57.Shetabivash H. Investigation of opening position and shape on the natural cross ventilation. Energy Build. 2015;93:1–15. doi: 10.1016/j.enbuild.2014.12.053. [DOI] [Google Scholar]
- 58.Etheridge D. first ed. John Wiley & Sons Ltd; 2012. Natural ventilation of buildings: theory, measurement and design. [DOI] [Google Scholar]
- 59.Edward E.J., Randall W.C. The neutral zone in ventilation. Trans. Am. Soc. Heat. Vent. Eng. 1926;32:59–74. [Google Scholar]
- 60.Andersen K.T. Optimal design and control of buoyancy-driven ventilation. Int. J. Vent. 2016;15:105–121. doi: 10.1080/14733315.2016.1203605. [DOI] [Google Scholar]
- 61.Ng K.C., Ng E.Y.K., Yusoff M.Z., Lim T.K. Applications of high-resolution schemes based on normalized variable formulation for 3D indoor airflow simulation. Int. J. Numer. Methods Eng. 2007;73:948–981. doi: 10.1002/nme.2106. [DOI] [Google Scholar]
- 62.ISO B. BSI; London: 2004. ISO 8996: 2004 Ergonomics of the thermal environment - determination of metabolic rate. [DOI] [Google Scholar]
- 63.Mukhtar A., Yusoff M.Z., Ng K.C. An empirical estimation of underground thermal performance for Malaysian climate. J. Phys. Conf. Ser. 2017;949:1–6. doi: 10.1088/1742-6596/949/1/012011. [DOI] [Google Scholar]
- 64.Kleijnen J.P. Springer; 2008. Design and analysis of simulation experiments. [DOI] [Google Scholar]
- 65.Pal P., Probst D., Pei Y., Zhang Y., Traver M., Cleary D., Som S. Numerical investigation of a gasoline-like fuel in a heavy-duty compression ignition engine using global sensitivity analysis. SAE Int. J. Fuels Lubr. 2017;10:56–68. doi: 10.4271/2017-01-0578. [DOI] [Google Scholar]
- 66.Wei T. A review of sensitivity analysis methods in building energy analysis. Renew. Sustain. Energy Rev. 2013;20:411–419. doi: 10.1016/j.rser.2012.12.014. [DOI] [Google Scholar]
- 67.Sofotasiou P., Calautit J.K., Hughes B.R., O'Connor D. Towards an integrated computational method to determine internal spaces for optimum environmental conditions. Comput. Fluids. 2016;127:146–160. doi: 10.1016/j.compfluid.2015.12.015. [DOI] [Google Scholar]
- 68.Kleijnen J.P.C., Sargent R.G. A methodology for fitting and validating metamodels in simulation. Eur. J. Oper. Res. 2000;120:14–29. doi: 10.1016/S0377-2217(98)00392-0. [DOI] [Google Scholar]
- 69.Kvalseth T.O. Note on the R2 measure of goodness of fit for nonlinear models. Bull. Psychonomic Soc. 1983;21:79–80. doi: 10.3758/BF03329960. [DOI] [Google Scholar]
- 70.Kleijnen J.P.C. Simulation optimization through regression or kriging metamodels. 2017. [DOI]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Data Availability Statement
Data will be made available on request.

















