Skip to main content
Scientific Reports logoLink to Scientific Reports
. 2025 Aug 29;15:31934. doi: 10.1038/s41598-025-17455-7

Machine learning approaches for predicting the construction time of drill-and-blast tunnels

Arsalan Mahmoodzadeh 4, Hamid Reza Nejati 1,, Nejib Ghazouani 2, Abdulaziz Alghamdi 3
PMCID: PMC12397348  PMID: 40883459

Abstract

This study examines the intricate task of predicting construction duration for drill-and-blast tunnels utilizing machine learning (ML) techniques. First, a comprehensive dataset (500 data points) encompassing 20 diverse parameters was compiled by constructing eight tunnels. After meticulous analysis, 17 of the 20 parameters were identified as crucial for training the algorithms. The overbreak and tunnel cross-section parameters were found to exert a significant influence on the tunnel construction duration. To enhance the predictive accuracy of the ML models, an intensive hyperparameter tuning process was conducted. The findings underscored the effectiveness of the Gaussian process regression model in capturing complex and nonlinear relationships, achieving an average R-squared of 0.89. Additionally, an ML-based graphical user interface (GUI) was developed to facilitate real‒time estimation of tunnel construction duration. This GUI not only enables initial predictions but also allows for dynamic updates throughout the construction phase, enhancing its practical utility.

Supplementary Information

The online version contains supplementary material available at 10.1038/s41598-025-17455-7.

Keywords: Drill-and-blast tunnels, Construction duration, Sensitivity analysis, Machine learning

Subject terms: Engineering, Mathematics and computing

Introduction

Building a tunnel is never a straightforward task; engineers regularly contend with hidden rock layers, shifting groundwater, extreme pressure, and increasingly stringent safety rules. Of all the worries that arise underground, the most stubborn is how long the work will take. Because the subsoil is blind to cameras and drilling tends to surprise everyone, delays can snowball, making it hard to stick to a budget or calendar1,2. Knowing the likely finish date lets managers spot trouble early, send crews and equipment at the right moment, and steer the project with less waste3. Fortunately, fresh digital tools (like virtual twins, machine learning (ML) forecasts, and sensors that relay data in real time) are starting to cut through the fog, allowing engineers to model risks on the fly and adjust well before the schedule slips4,5.

Recently, drill-and-blast tunneling has seen innovations that trim schedules, boost quality checks, and lift overall site safety6. Zhou et al.7refined blasting settings in hard-rock drives by adjusting overbreak and ground motion using a fractal metric, thereby improving both site security and wall finish. Zhang et al.8stacked meta-heuristic-guided SVMs into a single ensemble to forecast overbreak with narrow bands, providing teams with interpretable scores that mitigate schedule swings caused by hidden over-excavation. Shi et al.9 reviewed score cards for blast performance and rolled out a half-porosity index, stitched from 3D scans after each round, that clarifies how clean a blasted surface is. Rehman et al.10used finite-element modeling to investigate how detonation affects reinforced liners, revealing that blast frequency is as crucial to shell life as peak load. Working on Himalayan drives, Bajgain et al.11 listed site-level hurdles related to altitude and moisture, then outlined step-by-step overbreak cures that stabilize cycle times. Rounding out the set, Tian et al.12 tested wider-hole patterns in giant headings, showing that they can speed up advance and still meet above-ground quality targets—a win that could pull completion dates forward. Together, these studies highlight why today’s tunnel planners need duration models that rely on fresh data yet remain reliable when unexpected on-site issues arise. They also back the idea of pushing forward with cutting-edge ML tools to create forecasts that adapt quickly and calmly to shifting job conditions.

Artificial intelligence (AI) and ML have become transformative tools in engineering, enabling data-driven optimization, predictive maintenance, and real-time decision-making across disciplines such as structural analysis, geotechnical design, and construction management1316. Historically, the dearth of advancements in ML and AI methodologies has spurred a burgeoning interest in leveraging statistical and probabilistic techniques to enhance the accuracy and reliability of tunnel duration forecasting17,18. These methodologies harness historical data, empirical observations, and mathematical models to scrutinize the uncertainties and inherent variability in tunnel construction1921.

Many tunnel jobs achieve decent results with statistical and probabilistic techniques, especially when they start with clean, well-organized data, yet a set of deeper weaknesses still limits how widely and confidently those tools can be used. Perhaps the biggest hurdle is that the models break if the input numbers are off, even a small typo, a misread sensor, or an outlier rock sample can scatter the forecasts in surprising directions. Each tunnel then requires a fresh setup (custom equations, tailored samples, and site-specific distributions) before the model can do useful work, and that process consumes time, budget, and skilled hours. The underlying mathematics can also be intricate; building, testing, and fine-tuning a stable fit still requires solid training in statistics, alongside a feel for the ground, which not every project team can muster, especially under tight deadlines. On top of that, most studies rely on neat shortcuts (homogeneous layers, steady equipment speed, or constant advance rates) that real tunnels routinely disregard, undermining the trustworthiness of any results that rely on those simplifications. Finally, tunnel sites can move, the ground can shift, and machines can malfunction, so every prediction remains shrouded in a cloud of uncertainty that can grow alarmingly fast when hidden faults or weather delays emerge. Taken together, these pain points plead for smarter, agile tools that watch live project data, learn from it, and share the lessons rather than forcing every new drift back through the same brittle pipeline.

Given the inherent limitations and drawbacks of probabilistic and statistical methodologies, coupled with the burgeoning availability of data pertinent to tunnel construction timelines and the development of comprehensive tunnel databases, the foundation was incrementally established for the application of AI and ML techniques in estimating the time and cost of tunnel construction. ML models are often praised for their flexibility and ability to generalize, but they can easily stumble when trained on messy or incomplete data problem that regularly crops up in tunnel construction records. In response, this work establishes a rigorous preprocessing routine that identifies outliers, fills in missing values, standardizes scales, and verifies the distribution of each feature through multiple statistical checks. On the modelling side, ensemble algorithms such as Random Forest, which tend to shrug off small amounts of noise, are included alongside more sensitive techniques. Furthermore, cross-validation and sensitivity tests are conducted to determine whether a shift in the dataset still yields stable predictions. Taken together, these steps target data uncertainty head-on, boosting the confidence engineers can have in the estimated timelines of various tunnelling schemes.

Mahmoodzadeh et al.22 scrutinized the efficacy of Gaussian process regression (GPR) and support vector regression (SVR) in forecasting geological conditions, as well as the time and cost associated with the construction of road tunnel projects. They trained the ML algorithms using data from previously constructed road tunnels and observations from the tunnel under investigation. A comparative analysis of predictions from the GPR and SVR models against actual data showed that both methods produced acceptable and highly accurate forecasts. They continuously updated the models by integrating data from newly constructed sections into the initial training dataset, thereby significantly enhancing the accuracy of the ML models. However, these models exhibited a critical limitation by considering the Rock Mass Rating (RMR) parameter as the sole factor influencing construction time and cost, neglecting essential parameters such as the support system, drilling execution method, and tunnel geometry. This omission renders these models less applicable to other tunneling projects, introducing substantial uncertainty and risk into their estimates.

Mahmoodzadeh et al.23 enhanced the GPR algorithm by integrating it with meta-heuristic algorithms to estimate the time and cost of road tunneling projects. Initially, they amassed 900 data points encompassing 16 parameters that influence the time and cost of tunnel construction, based on historical tunnels. These parameters spanned geological conditions, support systems, drilling methods, and workforce composition (number of engineers and workers). Among these, the drilling method and support system emerged as the most significant factors influencing the time and cost of tunnel construction. Their analysis demonstrated that optimizing the GPR method with meta-heuristic algorithms substantially improved prediction accuracy. In a recent benchmarking study, Mahmoodzadeh et al.24 tested twelve ML algorithms to see which could best forecast both the duration and expense of drill-and-blast tunnel projects. Pulling records from thirteen completed tunnels and focusing on ten key variables, they put each model through the same rigorous checks. All twelve delivered solid predictions during training, yet only GPR kept that level of accuracy on fresh data, underscoring its remarkable ability to generalize beyond known cases.

ML algorithms have been developed, or are in the process of development, to a broadly acceptable extent in various aspects of tunneling. However, this advancement is markedly less evident in the domain of predicting tunnel construction times. The limited body of work in this area is fraught with numerous constraints and shortcomings. Consequently, given the paucity of studies employing ML algorithms for tunnel construction time prediction and the inherent limitations of these studies, substantial progress is still required. Earlier studies (e.g24,25. , have shown that ML techniques work well in tunnel projects. Building on that foundation, this paper presents a series of meaningful enhancements that distinguish it from the existing pipeline. To begin with, an expanded list of twenty input features, far more than previous studies included, draws data not only from geology and geometry but also from the actual support systems and excavation methods used on-site. In addition, the authors devised a careful, four-step feature-selection routine (correlation checks, P-value tests, sensitivity ranking with mutual information) which has not been applied so thoroughly in tunneling before. A fair performance contest of six algorithms, including GPR, SVR, K-nearest neighbors (KNN), decision tree regressor (DTR), artificial neural networks (ANN), and random forest (RF), now rests on 5-fold cross-validation, providing a more robust confidence estimate than earlier studies could offer. Real-world proof comes from a recent case at the Bakhan tunnel in Iran, where the tight integration of live excavation data allowed the model to refine itself and, in turn, increase accuracy by a notable margin. Lastly, a simple graphical user interface (GUI) provides site managers with little coding experience to run predictions directly, bridging a gap that often stalls research when it reaches the job site. The findings of this research can advocate for the integration of ML techniques into project management frameworks, offering a paradigm shift in how construction timelines are anticipated and managed.

Methodology

In this section, the entire research workflow is detailed, beginning with data collection and preprocessing, proceeding to model training and evaluation, and culminating in the deployment of the model as a decision-support system. The methodology is organized into six main steps: (1) cleansing and noise mitigation, (2) feature selection, (3) model selection and justification, (4) ML techniques, (5) evaluation and proposed deployment strategy, and (6) toward a decision support tool. The goal of this research is to evaluate and compare the predictive accuracy of six machine learning models for forecasting the construction duration of drill-and-blast tunnels and assessing their uncertainty robustness.

Data cleaning and noise mitigation techniques

Tunnel construction records are often messy; differences in recording habits, changing site conditions, and measurement errors result in a dataset that can vary significantly. To protect the models from this unpredictability, a careful preprocessing routine was implemented. The main steps were:

  • Outlier detection and removal: Anomalies that could skew learning were sought through the interquartile range (IQR) rule and z-score cut-offs, with suspect points either flagged or dropped.

  • Imputation of missing values: Gaps in entries were filled using KNN imputation, which borrows information from similar cases to ensure realistic replacements and minimize bias.

  • Normalization and scaling: Every continuous feature was centered and scaled using z-score standardization, a requirement for SVR, KNN, and ANN models that struggle when inputs span very different ranges.

  • Feature importance and noise sensitivity analysis: Random Forest scores and correlation charts highlighted the strongest predictors, allowing weaker, noisy variables to be set aside and improving model clarity.

  • Cross-validation and generalization testing: A standard 5-fold cross-validation was repeated five times to assess the model’s stability across different dataset splits. This approach tests whether the model still performs well when exposed to minor noise or specific groups of samples that behave differently.

Blending these cleaning steps with thorough cross-validation ensures that the final ML systems do not just show good numbers on a single train-test split, but also stand up to the usual messiness seen in actual tunnel construction data.

Feature selection

The following techniques were used to perform feature selection:

  • Pearson correlation coefficient: A Pearson correlation coefficient was calculated between each of the twenty input features and the target variable (construction time), using the full dataset of 500 samples. The resulting coefficients revealed which parameters had strong, linear links with the target. Variables that showed weak or even negative correlations were noted for review, but none were dropped yet, allowing the possibility that important, non-linear relationships still existed.

  • P-value analysis: A series of univariate linear regression models were fitted, one for each input feature against construction time, using the statsmodels library in Python. This procedure provided a P-value for every variable, indicating the likelihood that the variable was related to the target by chance. Features with P-values smaller than 0.05 were retained, while those above the threshold were flagged; the test was repeated on the entire dataset to maintain consistency.

  • Mutual information (MI) analysis: To identify non-linear dependencies, MI scores were calculated between each input feature and the target using the mutual_info_regression function in the scikit-learn (version 1.2) library, again over the full dataset. MI is a model-free measure of how much knowing one variable reduces uncertainty about the other, thereby adding value beyond Pearson and P-value tests by identifying hidden interactions. Features with low MI values were marked for later consideration but were only downgraded if no evidence from the other methods showed they still mattered.

Model selection and justification

This work compares six popular ML algorithms: GPR, SVR, KNN, DTR, ANN, and RF. They were selected because each employs a distinct learning approach (probabilistic, instance-based, tree-based, and deep) and all have a proven track record of handling nonlinear regression in complex, real-world data. Collectively, the models bring unique benefits:

  • GPR estimates the uncertainty of each prediction, a feature that is particularly important in construction, where budgets can shift rapidly.

  • SVR handles noisy outliers effectively and learns useful patterns even when the training samples are limited.

  • KNN provides a simple, local baseline that shows how much the verdict may change with new data.

  • DTR is transparent and mirrors the step-by-step decisions found in typical building processes.

  • ANN manages complicated, curved patterns and discovers hidden interactions without outside guidance.

  • RF merges multiple trees to hedge against bias and variance, and also identifies which features matter most, scoring high on both reliability and interpretability.

By examining these six models side by side, the research aims to determine which one (or which pair of them) best predicts the time it takes to build a tunnel, particularly when unexpected issues arise. Mixing straightforward and more advanced approaches enables researchers to balance accuracy against the ease of explaining the results, the computational time each model requires, and how well they tolerate minor errors in the input data. It also mirrors what project teams face in the field, as budget or staffing limits often prompt them to adopt a basic model when the situation demands it.

ML techniques

For benchmarking, the authors picked six well-known ML algorithms (GPR, SVR, ANN, RF, KNN, and DTR) that have performed reliably on structured engineering datasets. This lineup spans various learning styles, from the instance-based, non-parametric nature of KNN, through the ensemble power of RF, to the kernel techniques of GPR and SVR, culminating in the layered, deep-learning perspective offered by ANN. Each method brings strengths that align with the demands of tunnel-construction modeling. GPR stands out because it naturally provides a measure of uncertainty, an asset when observations are sparse or noisy. SVR and KNN act as simple, interpretable reference models that focus on local data behavior. RF and DTR retain clarity through feature-importance scores. At the same time, ANN is introduced to capture complex, high-dimensional patterns, assuming there is sufficient training data to support that depth of learning.

For readers seeking more detailed technical information about the algorithms, a comprehensive summary is provided in Appendix A.

Modeling and optimization

To achieve the best results while avoiding problems such as underfitting or overfitting, a careful round of hyperparameter tuning was conducted using a grid search and 5-fold cross-validation. For every model, a set of candidate values was pre-identified for the primary hyperparameters. For example, tree depth and minimum samples per leaf were varied for DTR and RF, the regularization weight and kernel type were adjusted for SVR, the neighbor count k was tested for KNN, and in the case of the ANN, the number of hidden layers, neurons in each layer, and learning rate were fine-tuned. Every possible setting was evaluated through the cross-validation folds, and the one yielding the highest average R² and the lowest mean absolute percentage error (MAPE) on those folds was selected. Following this structured tuning path allows each algorithm to run at its optimal performance, thereby enhancing the overall trustworthiness and consistency of the forecasts.

Toward a decision support tool

To make the model truly useful on the job site, the top-performing algorithms are packaged into a simple Python graphical interface. Project managers can get real-time duration predictions for any tunnel section and feed fresh measurements back into the tool as work continues. This feedback loop transforms the static forecast into a living guide, situated squarely between engineering theory and day-to-day project control.

Dataset preparation and exhaustive analysis

Dataset preparation

Data sits at the heart of every ML project; without a rich, well-organized dataset, even the most sophisticated models stumble and fail. That’s why the first step in any new system is to pinpoint the key variables that steer the outcome, and only then proceed to collect evidence that shows how each one behaves. For the present study, twenty features were compiled through discussions with tunnel engineers, a review of older papers on tunnel performance, and initial tests that assessed the strength of each factor’s association with the available data. The final list includes only items that the team could measure on-site, that appeared in records from several tunnel jobs, and that had already demonstrated their significance for both drilling speed and machine staffing. Editors deliberately spread the picks across four areas: (i) ground conditions, (ii) tunnel shape and size, (iii) design of supports and reinforcements, and (iv) the day-to-day running of the project. To ensure the dataset accurately reflects the challenges teams face underground, we compiled records from eight bored-and-blasted road tunnels constructed in various Iranian regions between 2012 and 2023. We reviewed field logs, design notes, geotechnical studies, and post-build as-built reports, resulting in a comprehensive dataset comprising 500 individual data points. The information covers both what was planned and what happened on-site, with each entry double-checked by site engineers. By way of illustration, lining thickness is recorded as the size the crew achieved once the work finished, not the figure given on the blueprints; shotcrete thickness, meanwhile, is broken into two clear pieces: the average layer on the tunnel walls (ShL), and the first protective cover on the face (SF), ensuring they do not get confused with each other. Keeping these measures apart lets us see better how early support from face shotcrete and later wall shotcrete together shape the speed at which the tunnel is driven forward. Additionally, the primary measure we are concerned with (construction time) is captured as the total man-hours spent for every meter of advance, averaged over a sliding window of tunnel length, so one day’s hiccups do not skew the picture. Researchers also examined how the measures interact with each other in the physical world. Take overbreak volume (GOB) and shotcrete thickness (ShL); when more rock is removed than planned, extra concrete is needed, so these two are linked. That connection stayed intact during preprocessing and was rechecked in the correlation review. General summary statistics for the dataset are listed in Tables 1 and 2, with Table 1 showing the continuous values and Table 2 covering the categorical data.

Table 1.

General specifications of numerical parameters.

Parameter Symbol Unit Mean Std Min Max
Cross-sectional area CA m2 90.02 12.20 50.62 100
Rock mass rating RMR 31.65 14.78 6.0 83.0
Geological strength index GSI 27.05 14.64 4.0 78.0
Spacing between IPEs IPES m 2.41 3.31 0.30 2.0
Mesh diameter MD mm 7.26 0.97 6.0 8.0
Mesh network MN cm2 144.5 59.91 100.0 225.0
Shotcrete thickness ShL cm 23.17 7.88 5.0 35.0
Shotcrete thickness on tunnel face SF cm 2.22 2.30 0.00 6.0
Lining thickness LT cm 36.83 5.79 25.0 45.0
Bolt length BL m 5.54 0.76 4.0 6.0
Bolt diameter BD mm 26.47 1.50 25.0 28.0
Bolt network BN m2 3.76 3.09 2.0 10.0
Explosive specific charge SC kg/m3 0.87 0.71 0.00 2.0
Forepoling Fo Percent of tunnel head circumference (%) 30.16 40.08 0.00 100.0
Drilling cycle DC m 1.65 0.95 0.40 4.0
Overbreak GOB m3 15.48 5.56 7.0 35.0
Construction time Hours per one meter of tunnel (h/m) 43.56 10.57 19.60 73.90

IPE (I-beam profiles in European standards).

Table 2.

General specifications of nominal parameters.

Parameter Symbol State Percent of total
Underground water GW Dry 0.14
Wet 0.32
Drop 0.24
Flow 0.30
Number of mesh layers NML One layer 0.27
Two layers 0.62
Three layers 0.11
IPE (I-beam profiles in European standards) IPE Not used 0.16
IPE16 0.12
IPE18 0.72
Number of drilling steps NDS Two-stages 0.67
Three-stages 0.17
Four-stages 0.16

Figure 1 uses violin plots to illustrate the distribution of all input and output variables, allowing readers to quickly identify each variable’s range, density, and number of peaks. Such a display helps identify skew, multiple clusters, and concentrated zones that could influence how well an ML model learns. The database behind these graphics was compiled from several real tunnel projects, and each entry was painstakingly checked to remove identical records, where every value matched across different tunnels. The variables shown in the plot combine both categorical and numerical data, and their shapes indicate the dataset’s variety and overall coverage. Wide and even spreads of numerical values usually help ML methods, because they mean no single condition dominates training. Take the RMR; its scores range from approximately 6 to 83, indicating that the records encompass tunnels dug through very soft to extremely tough rock. This broad range gives the models a solid chance to learn how shifts in RMR affect construction time under all the ground conditions represented by the data. Similarly, the thickness of each ShL and the amount of GOB follow skewed, non-normal distributions that mirror the everyday ups and downs found on site. These spreads matter because they help expose the nonlinear links analysts care about. In contrast, categorical items such as GW and NDS sit in clean, distinct buckets, allowing the models to tell one work situation from another without second-guessing. When viewed together, the plots do more than show whether any single variable dominates; they make it easy to see how balanced and varied the input is, and they reassure us that the data can withstand the broader test of generalization.

Fig. 1.

Fig. 1

Violin plots of inputs and output parameters.

Exhaustive analysis of the dataset

In the preceding section, we identified 20 parameters as influential on tunnel construction time. In this section, our objective is to diminish the dimensionality of the data matrix using statistical data analysis, thereby eliminating redundant parameters. Initially, we scrutinize the correlation between input parameters. If the correlation between two parameters is greater than or equal to 0.9, or less than or equal to -0.9, only one parameter is retained while the other is discarded. This is due to the fact that one parameter’s values can be accurately estimated from the other, rendering their effects on the output nearly identical.

The Pearson correlation coefficient is employed to quantify the correlation between input parameters. This coefficient, ranging from − 1 to + 1, measures the degree of relationship between two variables. A coefficient of + 1 indicates a perfect positive correlation, meaning that an increase or decrease in one variable corresponds proportionally with an increase or decrease in the other. Conversely, a coefficient of -1 signifies a perfect negative correlation, where a rise in one variable corresponds with a reduction in the other. A coefficient of zero denotes no relationship between the variables.

Figure 2 presents the correlation matrix derived via Pearson’s method among numerical parameters. It reveals that the correlation coefficient between the DC parameter and each of the RMR and GSI parameters is 0.93. Consequently, the DC parameter is eliminated. Furthermore, the correlation coefficient between the RMR and GSI parameters is 0.99, indicating that one parameter can be accurately inferred from the other. Hence, we remove the GSI parameter and retain the RMR parameter. The correlation coefficients between all other parameters are below 0.9 (and above − 0.9), necessitating the inclusion of these parameters. As a result, out of the initial 16 numerical input parameters, two parameters (GSI and DC) are excluded, leaving 14 parameters for further analysis.

Fig. 2.

Fig. 2

Pearson’s correlation between the numerical input parameters.

We now aim to evaluate the impact of each numerical input parameter on the tunnel construction time, eliminating those that do not significantly influence the output. This is achieved through parameter sensitivity analysis, a crucial method for discerning the most impactful inputs and discarding those with negligible effects. Incorporating superfluous variables adds unnecessary complexity to the model and can degrade its accuracy and execution efficiency. Moreover, reducing the dimensionality of input data simplifies the model, particularly for future applications involving new tunnels, by minimizing the number of required input parameters.

In this study, we employ the P-value method to identify redundant input variables. The P-value is a statistical measure used in feature selection to assess the significance of each variable within a model. It is commonly applied in statistical hypothesis testing, especially in the context of null hypothesis testing, where it is assumed that no relationship exists between the variables. The P-value helps determine the statistical importance of a feature relative to other variables. Suppose the P-value is below a specified significance threshold (typically 0.05). In that case, it indicates that the observed results are highly unlikely under the null hypothesis, allowing us to reject the null hypothesis and affirm that the feature has statistical significance. The P-values are typically calculated using a univariate linear regression model for each input variable with respect to the target variable. This classical statistical approach assesses the individual correlation and significance of each feature within a linear framework, serving as a preliminary screening method to identify features that have statistically meaningful relationships with the output.

It is essential to note that the P-value is calculated using different methods, depending on the type of parameters. Since our objective at this stage is to ascertain the correlation between numerical input parameters and construction time, and given that all parameters under consideration are numerical, the t-test method is employed to calculate the P-value.

The P-values reflecting the relationship between each numerical input parameter and the tunnel construction time are presented in Table 3. All numerical parameters, except the MD parameter, exhibit P-values less than 0.05. This indicates a significant relationship between these parameters and the construction time parameter, with a confidence level exceeding 95%. Therefore, except for the MD parameter, other parameters are considered influential on the construction time of tunnels.

Table 3.

The quantitative significance of each numerical parameter on the tunnel construction time was determined via P-value analysis.

Input parameter Output parameter P-value
CA Construction time Inline graphic
RMR Construction time Inline graphic
IPES Construction time 0.00075
MD Construction time 0.27183
MN Construction time Inline graphic
ShL Construction time Inline graphic
SF Construction time Inline graphic
LT Construction time Inline graphic
BN Construction time Inline graphic
BL Construction time Inline graphic
BD Construction time Inline graphic
SC Construction time Inline graphic
Fo Construction time Inline graphic
GOB Construction time Inline graphic

We now aim to examine the correlation among the non-numerical input parameters (GW, IPE, NML, and NDS) and the tunnel construction time parameter. For this purpose, we determine the correlation using the P-value. In this step, the Chi-square method is used to calculate the P-value. The resultant P-values for this scenario are presented in Table 4. The P-values for all these parameters in relation to the construction time parameter are below the threshold of 0.05. Consequently, these parameters are also recognized as significant non-numerical factors influencing the tunnel construction time.

Table 4.

The quantitative significance of each nominal parameter on the tunnel construction time was determined via P-value analysis.

Input parameter Output parameter P-value
GW Construction time Inline graphic
NDS Construction time Inline graphic
NML Construction time Inline graphic
IPE Construction time Inline graphic

The statistical analysis of the data revealed that, with the exception of the DC, GSI, and MD parameters, the remaining input parameters initially selected should be regarded as influential on tunnel construction time. Consequently, the dataset utilized in this research encompasses seventeen input parameters (GW, IPE, NML, DNS, CA, RMR, IPES, MN, ShL, SF, LT, BN, BL, BD, SC, Fo, and GOB) along with one output parameter (tunnel construction time). Figure 3 shown a visual flowchart summarizing the parameter selection process:

Fig. 3.

Fig. 3

Flowchart of feature selection process.

Encoding the nominal parameters

It is crucial to note that numerical and nominal parameters cannot be used concurrently within the same dataset for training ML models. To address this issue, it is essential to transform the nominal parameters into numerical codes. In this research, we employ the One-Hot encoding method. One-Hot encoding is a widely used technique, known for its effectiveness except when the categorical variable has an excessive number of values. Typically, this method is unsuitable for variables with more than 15 distinct values. One-Hot encoding generates new binary columns, each representing one of the variable’s possible values. To elucidate this process, we demonstrate the encoding of the GW parameter, which includes four states: Dry, Wet, Drop, and Flow. Suppose the GW parameter has five samples with the states Dry, Dry, Flow, Drop, and Wet, respectively, as shown in Fig. 4 (left side). Using One-Hot encoding, the values of this variable are transformed into four separate columns titled Dry, Flow, drop, and Wet, as illustrated in Fig. 4 (right side). Thus, instead of a single GW parameter, we have four distinct parameters named Dry, Flow, Drop, and Wet. For instance, in Fig. 4 (right side), the first sample is encoded with a 1 in the “Dry” column and 0 in the “Flow,” “Drop,” and “Wet” columns. This encoding is similarly applied to the second sample, which is also “Dry.” For the third sample, a 1 is entered in the “Flow” column and 0 in the remaining columns. This pattern is repeated for the other samples accordingly.

Fig. 4.

Fig. 4

An example of One-Hot encoding for the GW parameter.

Sensitivity analysis based on MI test

At an earlier stage of this study we used P-value tests to sift through the input data and see which factors really mattered for tunnel construction time. That round let us cut out noise and keep only variables that had a clear statistical link with the outcome. Yet P-values, useful as they are, fall short of showing how much each surviving factor pushes the time up or down, let alone putting them in a clear order of importance. To fill that gap, we now run a second round of checks, this time leaning on MI test, which looks at how much information each input actually shares with the output. Although we call both reviews sensitivity analysis, they serve different ends:

  • The P-value work in Sect. 3.2 zeroed in on whether a factor mattered at all;

  • The MI test here digs deeper, letting us score each factor by how strongly it moves construction time and stack them from most to least influential.

MI is a flexible statistical tool that measures how much knowing one variable reduces uncertainty about another, in this case between each input feature and construction duration. Because it captures both straight-line and twisting patterns, a higher MI score clearly shows stronger dependence between the two, no matter how complex the link. Using MI this way reveals whether geological conditions, structural choices, or work flows truly push the schedule longer or shorter.

The MI value is derived from Eq. 1, where Inline graphic and Inline graphic represent input and output parameters, respectively.

When holding the input parameter (Inline graphic) constant, the MI value becomes zero, indicating no discernible connection between parameter Inline graphic and the output (Inline graphic). The maximum MI value varies depending on the data type. If the input and output parameters possess identical or completely opposite values (i.e., an R2 value of 1 or -1), the MI value reaches its peak, signifying a linear relationship between the two parameters. Hence, the higher the MI value between two parameters, the more intertwined their relationship, facilitating the identification of one parameter’s state from the other.

graphic file with name d33e1266.gif 1

Initially, to calculate the maximum expected value of MI between an input parameter and the tunnel construction time, we set Inline graphic in Eq. 1. This means the tunnel construction time parameter is considered both as an input and an output. Under these conditions, the maximum expected MI value was determined to be 4.67. The closer the MI value between an input parameter and the tunnel construction time is to this maximum value, the more linear the relationship, and conversely, the further away, the less linear.

This study uses MI analysis for two main reasons. First, it gives a broad and sturdy measure of dependence that picks up both straight-line and twisty links between the input factors and construction time, capturing patterns traditional correlations or plain p-values might miss. That broad view matters when ML models like GPR and ANN are being trained, because those tools aim to learn the sort of complex, bent behaviour MI is built to find. Second, the MI scores let us rank the features by how much each adds to explaining the target outcome. That ranking not only makes the models easier to read, it can guide future work in thinning the dataset or picking the strongest predictors, especially when records are scarce. Although nothing was dropped solely because of a low MI score here, the exercise backed up earlier tests and shed light on how much each input really drives the models performance.

Table 5 presents the MI values between the input parameters and the tunnel construction time. According to this data, tunnel construction time exhibits the highest sensitivity to parameters related to the excavation method and geological conditions, such as GOB, CA, RMR, and IPES. Conversely, parameters associated with rockbolts, mesh, and IPE display the least sensitivity in predicting tunnel construction time. The lower sensitivity of these parameters can be attributed to their narrower range of variation within the dataset compared to other parameters.

Table 5.

MI scores between the input parameters and the tunnel construction time parameter.

GOB CA RMR NDS IPES Fo LT SF ShL SC GW NML BN BD BL MN IPE
0.53 0.53 0.45 0.44 0.43 0.36 0.33 0.32 0.31 0.31 0.21 0.20 0.19 0.14 0.11 0.08 0.07

Training and optimization of ML algorithms

Upon completing the data analysis and identifying the parameters significantly impacting tunnel construction time, the next step is to train the ML algorithms. It is imperative to meticulously select the type or value of algorithm parameters to avoid issues of underfitting and overfitting. Concurrently, the optimal value or type for each algorithm’s parameters must be considered to ensure the highest predictive accuracy. For instance, with the DTR algorithm, tree depth is a critical factor influencing overfitting. As illustrated in Fig. 5, the diagram’s horizontal axis represents tree depth, while the vertical axis indicates error magnitude. A lower error value signifies better model performance. In the case of the lower curve representing training data, increasing tree depth beyond a certain point (the optimal point) result in negligible changes in error value, indicating a progression toward overfitting. Conversely, the upper dashed curve, which pertains to test data, demonstrates that error decreases up to the optimal point but increases thereafter. Hence, the vertical line at the midpoint of the diagram indicates the appropriate tree depth for the DTR algorithm in this example, balancing complexity and predictive accuracy.

Fig. 5.

Fig. 5

How to detect the optimal value of the tree depth in the DTR algorithm to prevent over-fitting and under-fitting.

Overfitting is often more challenging to discern than underfitting because, unlike underfitting, an overfitted model achieves high accuracy on training data. One method to ascertain whether neither overfitting nor underfitting has occurred during algorithm training is to assess the algorithm performance on new data points and analyze the results using different statistical metrics. Additionally, techniques such as K-fold cross-validation are instrumental in evaluating algorithm results. All these three approaches have been employed in this research to rigorously evaluate the algorithms performance.

Results analysis and comparison

In this research, the 5-fold cross-validation method is employed to assess the performance of ML algorithms. Table 6 summarizes findings obtained through 5-fold cross-validation, a method that divvies the training set (%80 of the full data) into five separate mini-sets, allowing each one in turn to act as a temporary validation fold. By testing the model on data it has not yet seen within the training portion, this procedure gives a clearer picture of how well the model can generalize and helps guard against overfitting because results are averaged across all five splits. The score provided in Table 6 for each fold is a ranking index, not a raw metric. For each fold and each metric (R², MAPE, MSE), the six models were ranked from 1 (best) to 6 (worst). For R², a higher value corresponds to a better model, so ranks were assigned accordingly. For MAPE and MSE, a lower value indicates better accuracy, hence ranks were inverted. These rank scores were then summed across the three metrics for each fold, and a total ranking score was computed for each algorithm by summing all five folds. The final ranking score (last column) represents the cumulative performance of each algorithm across all folds and metrics. The higher the ranking score, the more consistently the algorithm outperformed others.

Table 6.

Ranking scores of the ML algorithms for predicting tunnel construction time using the 5-Fold cross-validation method.

Algorithm R 2 Score MAPE Score MSE Score Sum of scores Ranking score
DTR Fold 1 0.69 2 0.33 3 66.32 2 7 34
Fold 2 0.40 1 0.44 3 183.53 2 5
Fold 3 0.52 1 0.28 4 102.63 2 6
Fold 4 0.40 1 0.30 2 119.41 2 4
Fold 5 0.53 1 0.36 4 92.23 2 6
KNN Fold 1 0.56 1 0.36 2 108.0 1 4 31
Fold 2 0.68 2 0.48 2 89.91 2 6
Fold 3 0.65 2 0.35 2 78.39 2 6
Fold 4 0.67 2 0.29 3 66.86 2 7
Fold 5 0.58 2 0.46 1 83.75 2 5
SVR Fold 1 0.76 3 0.60 1 58.65 3 7 80
Fold 2 0.76 5 0.40 5 63.94 5 15
Fold 3 0.79 5 0.24 5 45.48 5 15
Fold 4 0.92 6 0.25 5 26.33 5 17
Fold 5 0.82 5 0.29 5 63.21 4 14
GPR Fold 1 0.90 6 0.21 6 30.22 6 18 89
Fold 2 0.87 6 0.16 6 29.51 6 18
Fold 3 0.89 6 0.22 6 37.11 6 18
Fold 4 0.91 5 0.21 6 24.18 6 17
Fold 5 0.88 6 0.20 6 41.10 6 18
RF Fold 1 0.82 5 0.26 5 43.59 5 15 60
Fold 2 0.74 4 0.42 4 72.88 4 12
Fold 3 0.73 4 0.32 3 60.40 4 11
Fold 4 0.74 3 0.27 4 51.43 3 10
Fold 5 0.70 4 0.41 3 59.16 5 12
ANN Fold 1 0.79 4 0.32 4 51.34 4 12 43
Fold 2 0.72 3 0.57 1 77.61 3 7
Fold 3 0.66 3 0.43 1 76.69 3 7
Fold 4 0.75 4 0.31 1 50.37 4 9
Fold 5 0.67 3 0.47 2 65.37 3 8

R2 (coefficient of determination); MAPE (mean absolute percentage error); MSE (mean squared error).

By analyzing the ranking scores obtained for each algorithm across different folds, it can be inferred that, except for the DTR and KNN algorithms, the others have yielded results with acceptable accuracy. An average R² value of 0.89 achieved by the GPR model indicates a very high level of predictive accuracy, suggesting that nearly 89% of the variance in tunnel construction time can be explained by the selected input parameters. In the context of tunneling projects (where time predictions are heavily influenced by complex, variable geological and operational factors), this level of accuracy is particularly noteworthy.

To compare the overall performance of the ML models, all scores obtained for each algorithm are aggregated. From the sum of these scores, a rank is assigned to each algorithm, as shown in the last column of Table 6. The GPR algorithm, with a ranking score of 89 (an average R2 of 0.89), demonstrates the highest accuracy and best performance. The remaining algorithms, in descending order of accuracy, are SVR, RF, ANN, DTR, and KNN.

Testing a ML model against an untouched 20% hold-out set is the only way to see how it will really perform on messy, noisy, and ever-changing tunnel sites in the real world. Consistent with that approach, Table 7 lists the final rankings of each algorithm using R², MAPE, and MSE, all calculated on that unseen data. Of the six methods we tried, GPR took the top spot once more, finishing with a combined score of 18 and proving it can generalize solidly across different project conditions. On the test set GPR also delivered the best R² of 0.90, the smallest MAPE at 0.19, and the lowest MSE of 30.17, leaving the other models behind on every front. By contrast, DTR posted a disappointing R² of only 0.43, a MAPE of 0.35, and an MSE that soared to 112.41, giving it a final ranking score of just 3. Those numbers underscore how single-tree approaches struggle to untangle the complex, non-linear patterns common in tunnel construction data. SVR and RF showed middling results compared to GPR. SVR scored a solid R² of 0.81, yet its higher error rates (MAPE = 0.32, and MSE = 55.04) kept it from topping the list. KNN and ANN looked decent too, but their ability to handle new data fell short of GPR, especially on the error fronts. Taken together, the tests prove that many ML techniques can be accurate, but only GPR delivered steady, trustworthy time estimates for tunnel projects when exposed to real-world cases unseen in training. This finding echoes what the training phase showed and underlines why looking at different error measures on a separate test set gives a fuller picture of how well a model will perform in practice.

Table 7.

Ranking scores of the ML algorithms for predicting tunnel construction time using hold-out cross-validation method.

Algorithm R 2 Score MAPE Score MSE Score Ranking score
DTR 0.43 1 0.35 1 112.41 1 3
KNN 0.62 2 0.42 2 98.59 2 6
SVR 0.81 5 0.32 5 55.04 5 15
GPR 0.90 6 0.19 6 30.17 6 18
RF 0.72 4 0.33 4 55.63 4 12
ANN 0.69 3 0.40 3 71.92 3 9

To judge whether feeding real-time data back into a ML model speeds up tunnelling work on the ground, a step-by-step test plan was designed that retrains the system after each small batch of excavation records comes in. In practice, the team adopts what is called batch learning, so the full dataset pauses until a set volume of new measurements arrives, rather than letting the model grow piece by piece (Fig. 6). During early trials the group picked GPR because it handled noisy subsurface signals well and delivered its uncertainty estimates fast. Updates trigger each time fresh data rolls in from the advancing face, mimicking site conditions where bore advance never stops for long. In this controlled scenario, the engineers adopted a straightforward milestone strategy: once approximately 200 m of spoil had been removed from both the tunnel entrance and exit, the corresponding datasets were merged, initiating the model retraining process.

Fig. 6.

Fig. 6

Model updating scheme for construction time prediction in drill-and-blast tunnel projects.

To see how well the approach works in practice, the team examined the 1848-meter Bakhan tunnel in Iran. At first, the GPR model, which relied only on past records, estimated the entire tunnels completion time with an R2 of 0.82 (see Fig. 7a). Later, data from two recently excavated stretches (250 m from the entrance and 200 m from the exit) were added and the model was retrained. The new forecast for the still-unexcavated segments, shown in Fig. 7b, tracked the real work much closer, lifting the R2 to 0.92. Together, these findings point to a clear benefit of routinely feeding field data back into the model to boost the accuracy of tunnel-time predictions.

Fig. 7.

Fig. 7

Predicted construction time for the 1848-meter-long Bakhan tunnel using the GPR model: (a) prior to updating; (b) post-update.

Graphical user interface (GUI)

Undoubtedly, the process of training ML algorithms and writing the corresponding code is an intricate and time-consuming endeavor, particularly for tunneling engineers who may lack proficiency in programming languages and the underlying algorithms. To address this challenge, this research has developed a graphical user interface (GUI) grounded in the pre-trained ML algorithms. The interface of this GUI is illustrated in Fig. 8. The GUI was built in Python 3.10 with the Tkinter toolkit, chosen for its small footprint, ease of learning, and broad cross-platform support. It works hand-in-hand with ML models created and stored using scikit-learns 1.2 release. When the window opens, anyone can pick a ML model and fill out simple form fields with project data. In the background the app loads the saved models with joblib or pickle, runs a quick one-hot encode, and returns an estimated construction time based on the numbers entered. Because the components are modular new models or features-batch prediction, graphing-drop in with little rework. This clean, tight setup lets tunnel engineers who dont code much tap cutting-edge ML insights on the job.

Fig. 8.

Fig. 8

The developed GUI predicated upon advanced ML algorithms for the precise prediction of construction time in drill-and-blast tunneling projects.

This GUI proves to be exceptionally valuable during the initial planning stages of tunneling projects. Moreover, it is an indispensable tool for individuals seeking to obtain new data concerning the construction time of drill-and-blast tunnels. The GUI’s utility is further enhanced by its capacity to generate extensive datasets through the application of the GPR algorithm, thereby facilitating future research endeavors. Consequently, this interface not only simplifies the computational complexities for engineers but also significantly contributes to the accumulation of pertinent data for ongoing and prospective studies in the field.

Key limitations and suggestions

Among the most significant limitations of this research, accompanied by corresponding recommendations, are the following:

  • In this research, only 500 data points were utilized to train and test the ML algorithms. Given the substantial number of input parameters considered, a larger dataset is required to ensure these algorithms are properly trained. The limited data in relation to the number of parameters is a principal reason why only the GPR algorithm consistently delivered accurate performance across all stages. It is recommended that future work expand the dataset to enhance the training and effectiveness of these algorithms.

  • All the data points employed in this study were derived from tunnels constructed using drilling-and-blasting method. With the continuous advancement of technology, tunnel excavation increasingly adopts mechanized methods. Consequently, estimating the construction time of these mechanized tunnels has become particularly significant. Identifying the parameters influencing the construction time of mechanized tunnels will enable the collection of pertinent data and the development ML models to predict their construction duration.

  • Although the ML models built so far predict tunnel-construction times using the data at hand, they leave out key quality-control and safety angles-like underbreak clean-up, shotcrete rebound and curing, dual-lane sprayer downtime, and extra ground support when geology turns bad. Veterans in the field already recognize these problems, and each one has been known to stall progress for days or weeks. The main reason the models ignore them is simple: the historical records that sit in the project database are neither uniform nor precise enough for reliable analysis. Moving forward, researchers should set up targeted logging (whether through check sheets, sensors, or digital diaries) so that those real-life disruptions show up as measurable inputs. By blending those fresh safety and quality signals into the existing datasets, the revised algorithm will give managers forecasts that align much better with the way work really unfolds underground.

  • The developed GUI possesses the capability to be updated concurrently with the construction of similar tunnels examined in this research. It is advisable to integrate new data into the GUI’s database to enhance its functionality and improve the accuracy of construction time estimates for the unconstructed sections of the tunnel.

  • Due to the limited available data, deep learning algorithms were not employed in this research. These algorithms excel in handling complex and large datasets. If a larger dataset were available, it is suggested to explore the performance of deep learning algorithms in estimating the construction time of tunnels.

Conclusions

This research investigated how ML methods can predict the time needed to complete drill-and-blast tunnels, drawing on real data from eight separate projects around Iran. The study led to several key findings:

  • Careful statistical tests including Pearson correlation, P-value checks, and MI scoring highlighted a short list of key inputs. GOB, CA, RMR, NDS, and IPES ranked as the most sensitive factors and shaped the models most strongly. These results back up the intuition that these elements drive on-site progress.

  • By trimming redundant features, the study made the models easier to read and faster to run with almost no loss of predictive power. Of the original twenty inputs, seventeen survived the culling while still matching the data closely.

  • Tuning the models and checking when they had really settled down followed a step-by-step plan that paired grid search with 5-fold cross-validation. That routine picked the best settings for each ML technique and kept the models from being too simple or too complex. Out of the ML algorithms used in this study, the GPR did best on every key measure (R², MAPE, and MSE) because it naturally handles uncertainty and bends to non-linear patterns, even with a small pile of data.

  • The GPR engine now guides tunnel-planning teams by turning site-specific and design-related data into reliable forecasts of how long construction will take. By sorting out which factors slow work or speed it up, engineers can fine-tune design choices, dig methods, or ground support to lift overall project pace.

  • The GPRs promise to learn and adapt was proven on the ground, shown in the 1848-meter Bakhan tunnel, a real project with deadlines and budgets on the line. When records from the first 450 m of progress were fed back into the model, its forecasts for the rest of the tunnel got noticeably sharper. That result proves the tool is worth loading early, yet turns even more valuable as a living dashboard that adjusts plans on the fly whenever new sensors or survey data come in.

  • To make the system easier for users, a simple GUI was built. Engineers can enter details about their project and get an instant, data-driven estimate of how long construction will take, all thanks to the pre-trained ML models running in the background. At the same time, the interface keeps capturing new input, which feeds back into the system so we can retrain the models and make them smarter over time.

In short, marrying these ML tools (especially GPR) with the day-to-day tunnel planning process gives teams a robust way to forecast and control schedule risk. Looking ahead, researchers should widen the database as much as possible, adding records from mechanized shields, mixed-face geology, and other scenarios that keep them out of the office. On top of that, pairing the models with deep-learning layers and streaming sensor feeds could sharpen predictions even further and bring a full digital twin to life underground.

Supplementary Information

Below is the link to the electronic supplementary material.

Supplementary Material 1 (107.6KB, pdf)

Acknowledgements

The authors extend their appreciation to the Deanship of Scientific Research at Northern Border University, Arar, KSA for funding this research work through the project number “NBU-FFMRA-2024-2105-12“.

Author contributions

Arsalan Mahmoodzadeh: Data curation, Conceptualization, Investigation, Methodology, Resources, Writing – original draftHamid Reza Nejati: Investigation, Validation, Visualization, Writing - Review and EditingNejib Ghazouani: Visualization, Writing – original draftAbdulaziz Alghamdi: Writing - Review and Editing.

Funding

The authors extend their appreciation to the Deanship of Scientific Research at Northern Border University, Arar, KSA for funding this research work through the project number "NBU-FFMRA-2024-2105-12".

Data availability

Data not available due to restrictions imposed by research sponsors, ongoing analysis for future studies, and the necessity to maintain data confidentiality until further validation and publication. However, are available from the corresponding author on reasonable request.

Declarations

Competing interests

The authors declare no competing interests.

Footnotes

Publisher’s note

Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

References

  • 1.Sharafat, A., Latif, K., Park, S. & Seo, J. Digital Twin-Driven Optimization of Blast Design for Underground Construction, in Korean Society of Civil Engineers Conference, pp. 1–2. (2022).
  • 2.Sharafat, A., Latif, K. & Seo, J. Risk analysis of TBM tunneling projects based on generic bow-tie risk analysis approach in difficult ground conditions. Tunn. Undergr. Space Technol.111, 103860. 10.1016/j.tust.2021.103860 (May 2021).
  • 3.Sharafat, A., Tanoli, W. A., Zubair, M. U. & Mazher, K. M. Digital Twin-Driven Stability Optimization Framework for Large Underground Caverns, Applied Sciences, vol. 15, no. 8, p. 4481, Apr. (2025). 10.3390/app15084481
  • 4.Zhang, K., Zhang, A., Wang, X. & Li, W. Deep-learning-based point cloud completion methods: A review. Graph. Models. 136, 101233. 10.1016/j.gmod.2024.101233 (Dec. 2024).
  • 5.Hartl, I., Sorger, M., Hartl, K., Ralph, B. J. & Schlögel, I. Passive seismic monitoring in conventional tunnelling – An innovative approach for automatic process recognition using support vector machines, Tunnelling and Underground Space Technology, vol. 137, p. 105149, Jul. (2023). 10.1016/j.tust.2023.105149
  • 6.Periku, E. & Aga, A. Construction time analysis for different steps in Drill-And-Blast method of hydro power tunnel excavation. Int. J. Eng. Res. Appl.5 (1), 95–101 (2015). [Google Scholar]
  • 7.Zhou, J., Gao, S., Luo, P., Fan, J. & Zhao, C. Optimization of blasting parameters considering both vibration reduction and profile control: A case study in a mountain hard rock tunnel. Buildings14 (5), 1421. 10.3390/buildings14051421 (May 2024).
  • 8.Zhang, Y. et al. Advancing overbreak prediction in drilling and blasting tunnel using MVO, SSA and HHO-based SVM models with interpretability analysis, Geomechanics and Geophysics for Geo-Energy and Geo-Resources, vol. 11, no. 1, p. 53, Dec. (2025). 10.1007/s40948-025-00963-1
  • 9.Shi, J., Wang, Y., Yang, Z., Shan, W. & An, H. Comprehensive review of tunnel blasting evaluation techniques and innovative half porosity assessment using 3D image reconstruction. Appl. Sci.14 (21), 9791. 10.3390/app14219791 (Oct. 2024).
  • 10.Rehman, J. U., Park, D. & Ahn, J. K. Predicting Blast-Induced damage and dynamic response of Drill-and-Blast tunnel using Three-Dimensional finite element analysis. Appl. Sci.14 (14), 6152. 10.3390/app14146152 (Jul. 2024).
  • 11.Bajgain, S., Rawat, B. & Wagle, M. Enhancing tunnel construction efficiency in nepal: challenges and Over-Break mitigation in Drill-and-Blast tunneling. IJMIR Journal, 1, 2, doi: (2024). https://www.ijmir.com/v1i2/8.php
  • 12.Tian, X., Tao, T. & Xie, C. Research on the theory and method of reduced-hole blasting for large cross-section tunnel based on explosive energy dissipation, Geomechanics and Geophysics for Geo-Energy and Geo-Resources, vol. 10, no. 1, p. 96, Dec. (2024). 10.1007/s40948-024-00816-3
  • 13.Deng, T., Sharafat, A., Lee, S. & Seo, J. Automatic Vision-Based Dump Truck Productivity Measurement Based on Deep-Learning Illumination Enhancement for Low-Visibility Harsh Construction Environment, Journal of Construction Engineering and Management, vol. 150, no. 11, Nov. (2024). 10.1061/JCEMD4.COENG-14194
  • 14.Latif, K., Sharafat, A. & Seo, J. Digital Twin-Driven Framework for TBM Performance Prediction, Visualization, and Monitoring through Machine Learning, Applied Sciences, vol. 13, no. 20, p. 11435, Oct. (2023). 10.3390/app132011435
  • 15.Chimunhu, P., Topal, E., Ajak, A. D. & Asad, W. A review of machine learning applications for underground mine planning and scheduling. Resour. Policy. 77, 102693. 10.1016/j.resourpol.2022.102693 (Aug. 2022).
  • 16.Zhang, Y. et al. A visual survey of tunnel boring machine (TBM) performance in tunneling excavation: mainstream direction, brief review and future prospects. Appl. Sci.14 (11), 4512. 10.3390/app14114512 (May 2024).
  • 17.Mahmoodzadeh, A. et al. A Markov-based prediction model of tunnel geology, construction time, and construction costs, Geomechanics and Engineering, vol. 28, no. 4, p. Geomechanics and Engineering Volume 28, Number 4, (2022). 10.12989/gae.2022.28.4.421
  • 18.Mahmoodzadeh, A. et al. Predicting construction time and cost of tunnels using Markov chain model considering opinions of experts. Tunn. Undergr. Space Technol.116, 104109. 10.1016/j.tust.2021.104109 (Oct. 2021).
  • 19.Sinfield, J. V. & Einstein, H. H. Evaluation of tunneling technology using the ‘decision aids for tunneling,’ Tunnelling and Underground Space Technology, vol. 11, no. 4, pp. 491–504, Oct. (1996). 10.1016/S0886-7798(96)89245-5
  • 20.Einstein, H. H. Decision Aids for Tunneling: Update, Transportation Research Record: Journal of the Transportation Research Board, vol. 1892, no. 1, pp. 199–207, Jan. (2004). 10.3141/1892-21
  • 21.Einstein, H. H., Indermitte, C., Sinfield, J., Descoeudres, F. P. & Dudt, J. P. Decision aids for tunneling. Transp. Res. Record: J. Transp. Res. Board.1656 (1), 6–13. 10.3141/1656-02 (Jan. 1999).
  • 22.Mahmoodzadeh, A. et al. Decision-making in tunneling using artificial intelligence tools. Tunn. Undergr. Space Technol.103, 103514. 10.1016/j.tust.2020.103514 (Sep. 2020).
  • 23.Mahmoodzadeh, A. et al. Developing six hybrid machine learning models based on Gaussian process regression and meta-heuristic optimization algorithms for prediction of duration and cost of road tunnels construction. Tunn. Undergr. Space Technol.130, 104759. 10.1016/j.tust.2022.104759 (Dec. 2022).
  • 24.Mahmoodzadeh, A. et al. A rigorous examination of twelve Cutting-Edge Machine-Learning techniques for predicting time and cost in tunneling projects. J. Constr. Eng. Manag.150 (12). 10.1061/JCEMD4.COENG-15070 (Dec. 2024).
  • 25.Mahmoodzadeh, A. et al. Forecasting tunnel geology, construction time and costs using machine learning methods. Neural Comput. Appl.33 (1), 321–348. 10.1007/s00521-020-05006-2 (Jan. 2021).

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Supplementary Material 1 (107.6KB, pdf)

Data Availability Statement

Data not available due to restrictions imposed by research sponsors, ongoing analysis for future studies, and the necessity to maintain data confidentiality until further validation and publication. However, are available from the corresponding author on reasonable request.


Articles from Scientific Reports are provided here courtesy of Nature Publishing Group

RESOURCES