Skip to main content
ACS Omega logoLink to ACS Omega
. 2023 Feb 11;8(7):7201–7210. doi: 10.1021/acsomega.3c00289

Productivity Index Prediction for Single-Lateral and Multilateral Oil Horizontal Wells Using Machine Learning Techniques

Omar Q Alharbi 1, Sulaiman A Alarifi 1,*
PMCID: PMC9947949  PMID: 36844581

Abstract

graphic file with name ao3c00289_0008.jpg

Horizontal wells are geometrically shaped differently and projected to different flow regimes than vertical wells. Therefore, the existing laws that govern flow and productivity in vertical wells are not applicable to horizontal wells directly. The objective of this paper is to develop machine learning models that predict well productivity index using several reservoir and well inputs. Six models were developed using the actual well rate data from several wells divided into single-lateral wells, multilateral wells, and a combination of single-lateral and multilateral wells. The models are generated using artificial neural networks and fuzzy logic. The inputs used to create the models are the typical inputs used in the correlations and are well-known for any well under production. The results of the established ML models were excellent as suggested by an error analysis performed, reflecting the models to be robust. The error analysis showed high correlation coefficient values (between 0.94 and 0.95) supported by a low estimation error for four models out of six. The added value of this study is the developed general and accurate PI estimation model that overcomes many limitations of several widely used correlations in the industry and can be utilized for single-lateral or multilateral wells.

1. Introduction

Productivity index (PI) is an essential element of the oil and gas well performance evaluation package. It is defined as the quantity of oil produced for a certain drop in pressure. Alternatively, PI is the production flow rate (q) over the drawdown pressure (Δp), which can be expressed as in eq 1. PI unit measurements in field units is bbl/day/psi.

1. 1

Using Darcy’s definition for single-phase flow in a radial system under steady-state conditions (eq 2), PI for vertical wells can be re-expressed as in eq 3.1

1. 2
1. 3

Also, eq 5 expresses the PI for pseudosteady state (PSS) conditions using Darcy’s definition under PSS conditions (eq 4).2

1. 4
1. 5

The productivity index is primarily used to evaluate the performance of a certain well, which indeed helps to decide the well’s economic viability. However, this critical parameter is more difficult to obtain for horizontal wells. This is mainly attributed to the different flow behavior observed in horizontal wells, as well as different geometrical factors affecting the well’s productivity.

Generally, horizontal wells perform better than vertical wells because of the larger flow area that horizontal wells can cover. In addition, horizontal wells can be very effective in scenarios where vertical wells would be considered unfavorable, e.g., producing from a thin reservoir. Moreover, an accurate well performance estimate is essential. Therefore, multiple studies attempting to develop PI correlations for horizontal wells are found in the literature. The following table (Table 1) summarizes the different widely used correlations to estimate the PI for horizontal wells and their main assumptions.

Table 1. Summary of Different Widely Used Correlations to Estimate PI for Horizontal Wells.

correlation equation main assumptions
Borisov3
graphic file with name ao3c00289_m011.jpg 6
isotropic reservoir
steady-state regime
circular drainage area
Giger et al.4
graphic file with name ao3c00289_m012.jpg 7
isotropic reservoir
steady-state regime
elliptical geometry
Joshi5
graphic file with name ao3c00289_m013.jpg 8
anisotropy
steady-state regime
elliptical geometry
Babu and Odeh6
graphic file with name ao3c00289_m014.jpg 9
anisotropy
pseudosteady state flow
box-shaped reservoir
damage near wellbore
Economides et al.7
graphic file with name ao3c00289_m015.jpg 10
anisotropy
steady-state flow
elliptical drainage area
Renard and Dupuy8
graphic file with name ao3c00289_m016.jpg 11
steady-state flow
circular, elliptical, or rectangular drainage area
Butler9
graphic file with name ao3c00289_m017.jpg 12
anisotropy
box-shaped reservoir
Permadi10
graphic file with name ao3c00289_m018.jpg 13
anisotropy
steady or pseudosteady state
elliptical drainage area
damage near wellbore
Elgaghah et al.11
graphic file with name ao3c00289_m019.jpg 14
splitting up the drainage area into three parts. One semicircle part and two rectangle parts with different lengths
Hongen12
graphic file with name ao3c00289_m020.jpg 15
steady state
treated the horizontal well as a vertical well
Lu13
graphic file with name ao3c00289_m021.jpg 16
anisotropy
elliptical shape
Furui et al.14
graphic file with name ao3c00289_m022.jpg 17
anisotropy
radial flow near the wellbore and linear flow far away from the wellbore
damage near wellbore
Escobar et al.15
graphic file with name ao3c00289_m023.jpg 18
steady state
based on Joshi’s and Furui’s models
Farman and Abdulamir16
graphic file with name ao3c00289_m024.jpg 19
based on Economides’s model
damage near wellbore

Alternative approaches to conventional methods were covered in the literature, including utilizing machine learning (ML) techniques. Alarifi et al. introduced several models that employ machine learning methods, e.g., artificial neural network (ANN), fuzzy logic (FL), and functional networks (FN), to model the productivity index of single-lateral horizontal wells. Moreover, Alarifi et al. proceeded to compare the developed ML models with a well-known correlation, e.g., Furui and Joshi, demonstrating an outperformance of ML models over these correlations. Table 2 illustrates a comparison between the error, root mean square error (RMSE), correlation coefficient (CC), average absolute error (AAE), and maximum and minimum absolute errors (AE), obtained from the PI models of several well-known correlations in the industry, and the error obtained from the three developed machine learning models.17

Table 2. Comparison between ML vs Well-Known Correlations17.

model Joshi and Economides Borisov Giger–Reiss–Jourdan Renard–Dupuy Butler Furui ANN FL FN1 FN2
RMSE 43.42 156.78 63.91 89.20 66.92 66.61 22.21 18.83 23.77 23.85
CC 0.79 0.45 0.46 0.47 0.46 0.71 0.94 0.90 0.83 0.82
AAE 28.93 76.08 45.42 49.54 47.92 46.51 16.75 13.83 18.66 18.87
maximum AE 184.02 971.90 295.13 516.53 296.87 296.5 57.21 56.63 56.16 55.88
minimum AE 0.13 0.37 0.07 0.64 0.21 0.19 0.02 0.04 0.50 0.48

Al-Mashhad et al. proposed a machine learning model to evaluate the productivity index of multilateral oil horizontal wells. The ML approach used to develop the model was ANN. As a way to prove the robustness of the developed ML model, Al-Mashhad et al. compared the results with Borisov’s correlation. The results showed superior performance of the ML model over Borisov’s results, as demonstrated in Table 3, which includes average percentage relative error (Er), average absolute percentage relative error (Ea) and its minimum and maximum, spread of the data around the mean (S), and correlation coefficient (R or CC).18

Table 3. Comparison between the ANN Model and Borisov’s Correlation18.

  predicted from ANN
 
measurement training test overall calculated from Borisov
average percentage relative error –0.48 –2.93 –1.17 30
average absolute percentage relative error 5.89 12.83 7.85 50.4
minimum absolute percentage relative error 0.04 0.07 0.04 3.4
maximum absolute percentage relative error 38 39 39 143
spread of the data around the mean 9.4 17.1 12.1 53.8
CC 0.986 0.914 0.968 0.3

Horizontal wells have proved their dominance in the production field over their counterparts, vertical wells. However, horizontal wells are geometrically shaped differently and projected to different flow regimes than what is observed in vertical wells. Therefore, the existing laws that govern the flow and productivity in vertical wells are not applicable to horizontal wells, at least not directly. The literature covers numerous research papers that attempted to find alternative approaches to assess the performance of horizontal wells using the productivity index. Some of these correlations are very well-known, widely accepted, and applied in the industry, such as Joshi’s, Babu, and Odeh and Furui’s correlations. Recently, ML has paved its way through the petroleum industry. This remarkable tool has proven its competence to solve complex modeling problems. It is widely used to predict industry-related parameters, especially those that conventional methods could not solve or solved with high uncertainty, e.g., permeability prediction. This paper will approach the productivity index prediction problem using multiple ML techniques.

The objective of this paper is to employ machine learning techniques to develop a robust model to estimate the PI of horizontal oil wells accurately. The developed model is planned to achieve comparable, if not better, results than the existing well-known correlations, such as Joshi’s and Furui’s. Additionally, some of the developed models used the productivity index estimate calculated using well-known correlations, namely, Borisov, Joshi, and Furui, as inputs in the machine learning (ML) models. Also, to develop a correlation of the resulting weight and bias matrices, which can be then utilized for different data sets of the same input type. The added value of this study is the developed general and accurate PI estimation model that overcomes the many limitations of several widely used correlations in the industry. However, it needs to be taken into account that the ML models are developed for specific ranges of input data sets.

2. Data Preparation

The data under investigation is of two main categories: production data of single-lateral wells and production data of multilateral wells. The data collected for this work is from a valid rate test of several producing horizontal oil wells from a reservoir in the Middle East. All of the wells are currently producing above bubble point pressure. The data consists of 136 different laterals from 28 multilateral wells and 110 rate test data from 110 single-lateral wells (Table 4). Models were developed for each category, including an additional category that includes the combined data of both.

Table 4. Number of Data Points for Each Model.

data set single-lateral multilateral combined
number of data points 110 136 246

For single-lateral production data, the available parameters from the 110 wells are reservoir thickness (h), average permeability (k), effective horizontal length (Le), pressure difference between reservoir, wellbore (ΔP), skin factor (s), and oil productivity index (PI). Additional inputs were added as combinations of the provided parameters. For example, PI values using Joshi’s, Furui’s, and Borisov’s correlations were obtained using other inputs, as well as some reservoir fluid, rock, and geometry constants, e.g., viscosity (μo), oil formation volume factor (Bo), anisotropy constant (β), and wellbore radius (w). Additional transformations of the existing inputs were tested as well, such as log(k), ΔP·k, ΔP·log(k), and ΔP/Le.

Statistical analysis was prepared for the inputs to examine each input and investigate if a relationship exists with the output (PI) using the correlation coefficient (CC). Also, another objective was to remove any outlier data that could distort the model development process. A statistical summary of the data is presented in Tables 5 and 6, where Table 5 shows the raw data and Table 6 shows the inputs of transformation and combination.

Table 5. Statistical Summary of Single-Lateral Well Data.

parameter h (ft) k (mD) Le (ft) ΔP (psi) skin PI (bbl/day/psi)
minimum 5 2 165 2 –7.4 1.9
maximum 256 5626 3600 1043 180.0 385.7
mean 146 632 1175 193 12.6 65.0
standard deviation 53 741 746 220 30.8 66.1
correlation coefficient with PI 0.15 0.32 0.39 –0.52 –0.26 1.00

Table 6. Statistical Analysis of Single-Lateral Well Transformed or Combined Data.

parameter log(k) (mD) ΔP·k (psi mD) ΔP·log(k) (psi mD) ΔP/Le (psi/ft) Borisov (bbl/day/psi) Joshi (bbl/day/psi) Furui (bbl/day/psi)
minimum 0.30 6.55 × 102 6.8 0.001 0.2 0.1 0.1
maximum 3.75 1.21 × 106 2705.5 2.91 225.9 357.1 122.9
mean 2.60 9.54 × 104 475.4 0.29 40.8 49.3 17.6
standard deviation 0.50 1.77 × 105 545.2 0.46 41.3 63.9 23.0
correlation coefficient with PI 0.44 –0.22 –0.48 –0.43 0.75 0.77 0.71

Production data of the 28 multilateral wells with 136 different laterals include the average permeability (k), effective horizontal length (Le), wellbore pressure (Pwf), reservoir pressure (Pr), pressure difference between reservoir and wellbore (ΔP), percentage opened of the choke (%choke), and oil productivity index (PI). Additional transformations were added, such as log(k) and ΔP/Le. Table 7 summarizes the statistical analysis of the multilateral well production data.

Table 7. Statistical Summary of Multilateral Well Data.

parameter k (mD) log(k) (mD) Le (ft) ΔP (psi) %choke (%) Pwf (psi) Pr (psi) ΔP/Le (psi/ft) PI (bbl/day/psi)
minimum 3 0.48 2794 662 5 227 1831 0.09 0.7
maximum 253 2.40 12 335 1948 100 1621 2754 0.56 8.2
mean 22 1.22 6951 1325 36 919 2244 0.22 3.6
standard deviation 26 0.31 2327 251 19 242 209 0.10 1.6
correlation coefficient with PI 0.06 0.22 0.09 –0.01 0.44 0.001 –0.01 –0.14 1.00

To generalize the model, the data of the single-lateral and multilateral wells were combined. However, these two categories do not include the same type of data. Therefore, only data available in both data sets are used. These are average permeability (k), effective horizontal length (Le), pressure difference between reservoir and wellbore (ΔP), and oil productivity index (PI). Once again, further transformations are tested including log(k), ΔP·k, ΔP·log(k), and ΔP/Le.

This combined data set is then tested statistically for the input–output relationship, which would be useful in the model development stage. Table 8 sums up important statistical parameters of the data, such as minimum, maximum, average, standard deviation, and correlation coefficient between input and output.

Table 8. Statistical Summary of the Combined Data.

parameter k (mD) log(k) (mD) Le (ft) ΔP (psi) ΔP·k (psi mD) ΔP·log(k) (psi mD) ΔP/Le (psi/ft) PI (bbl/day/psi)
minimum 2 0.30 165 2 6.55 × 102 7 0.001 0.7
maximum 5626 3.75 12 335 1948 1.21 × 106 3148 5.95 385.7
mean 295 1.84 4369 819 5.91 × 104 1111 0.93 31.1
standard deviation 580 0.79 3393 612 1.25 × 105 780 0.96 53.7
correlation coefficient with PI 0.52 0.64 –0.44 –0.63 –0.02 –0.60 –0.49 1.00

3. Model Development

Two machine learning techniques (fuzzy logic and artificial neural networks) are used to estimate the productivity index of hundreds of wells. Fuzzy logic applications have significantly increased over the past recent years. This can be attributed to the various applications where this powerful tool can be utilized. This logical system, fuzzy logic, can be defined as “an extension of multivalued logic”. However, in a wider sense, fuzzy logic (FL) is almost synonymous with the theory of fuzzy sets, a theory related to classes of objects with unsharp boundaries in which membership is a matter of degree. This notability, among other AI techniques, comes from many sources, such as its flexibility and easy-to-understand concept.19 FL is known to help deal with uncertainty in engineering. It is an efficient technique to solve complex problems that have high uncertainty and fuzziness of information. It will allow us to handle imprecise knowledge and provide a powerful framework for reasoning.

The FL method selected in this paper for modeling is subtractive clustering. The model is built using an adaptive neurofuzzy inference system, which is a hybrid learning algorithm to identify the membership function parameters of single output, Sugeno-type fuzzy inference systems (FIS). A combination of least-squares and back-propagation gradient descent methods are used for training FIS membership function parameters to model the given set of input/output data. This fast method is best used when there is no clear idea about the data set cluster number. It approximates the data set cluster number and cluster centers, in which later these estimated elements are used to optimize the clustering and model identification process. Ultimately, subtractive clustering generates a fuzzy inference system that describes the input–output behavior.19

Inspired by the biological nervous system, an artificial neural network (ANN) is a system consisting of simple elements (neurons). The network can be trained using weights as a connecting factor between elements to perform a certain function. Based on the connection between elements, different network algorithms can be selected. This robust AI tool has proven its strength in many applications, such as classification, pattern recognition, and identification.

Basically, ANN consists of an input layer, an output layer, and hidden or more layers. The input layer is connected to a hidden layer, and the hidden layer is connected to the next hidden layer if available, or if there is only one hidden layer, it is connected to the output layer. The connection between the layers is achieved through weights and biases. In a typical neural network, P refers to the input, R is the number of inputs, w is the weight connecting between layers (e.g., input-to-neuron connection), b is the neuron bias, and “n” is assigned as “Wp + b” and inserted in “f”, which represents a transfer function. The objective of a transfer function is to take “n” and produce an output “a”, which will be assigned differently to each available neuron. The produced values of “a” will be ultimately used to compute the output.20

This work employs the back-propagation neural network (BPNN) technique in the modeling process. In simple terms, BPNN is based on generalizing the Widrow–Hoff learning rules. It follows a gradient descent algorithm in which it moves the weights along the negative of the gradient performance function to achieve a particular solution. The network architecture utilized was feedforward. It is a common network architecture used with a back-propagation algorithm.20

The process of model development consisted of six models, in which two models were developed for each data set (single-lateral, multilateral, and combined). Table 9 sums up the ML methods used for each model. These models were developed using the aforementioned ML techniques, FL and ANN. For each model, the data was split into 70% for training and 30% for testing for fuzzy logic models and 70% for training, 15% for internal validation, and 15% for testing for ANN models.

Table 9. Summary of the ML Methods Used for Each Model.

data set single-lateral multilateral combined data set
model number 1 2 3 4 5 6
ML method FL ANN FL ANN FL ANN

To assess a model’s quality, several parameters can be utilized. A correlation coefficient (CC) is a statistical parameter that tests the linear relationship between actual values and predicted values. Statistically, CC values can be expressed in the range of −1 to 1, with values closer to 1 indicating a strong positive relationship, while values approaching −1 represent a negative relation. In terms of the use of CC in model assessment, the closer the value of CC is to 1, the lower the error between actual and predicted values. Another parameter used to judge the goodness of a model is root mean square error (RMSE). RMSE is the average squared difference between the actual values and their corresponding predicted values and can be described mathematically as follows (eq 20)

3. 20

where n is the number of points in the data set. Lower values of RMSE are desired for better models. An additional parameter that was employed to evaluate the model is the average absolute percentage error (AAPE). AAPE is the average of the percentage difference between actual and predicted values divided by their actual values. To have a better model, the goal is to aim for lower values of AAPE. AAPE can be expressed as (eq 21)

3. 21

Hundreds of trials were run, including changing the model inputs used, until an optimum model is found for each data set (lowest prediction error). The final inputs that were used to give the best results for each model are listed in Table 10.

Table 10. Final List of Inputs Used for Each Model.

model inputs
1 h S Le ΔP log k   ΔP·log(k) ΔP/Le Joshi Furui Borisov
2 h S Le ΔP log k   ΔP·log(k) ΔP/Le Joshi    
3 k   Le ΔP log k       Pwf Pr %choke
4 k   Le ΔP       ΔP/Le Pwf Pr %choke
5     Le ΔP log k   ΔP·log(k) ΔP/Le      
6     Le ΔP log k ΔP·k ΔP·log(k) ΔP/Le      

The use of transformed parameters, such as an input that is a product of two others, as model inputs helped in the modeling process, providing better results. Also, using alternative correlations (Borisov, Joshi, and Furui) in the development of the ML models as part of the inputs showed improved results. Such utilization of these correlations, which considers flow regimes and geometrical factors, would add further value and better accuracy to the developed models.

The fuzzy logic systems used to develop the model were Sugeno fuzzy systems, specifically, using a subtractive clustering algorithm. The process of this method starts with the calculation of the likelihood for each data point to define a cluster center according to the surrounding data point density. Second, the data point with the highest potential to be the first cluster center is chosen and the remaining data points surrounding it are removed. Then, the next data point with the highest potential from among the remaining data points is chosen and these steps are repeated until all of the data is within a cluster center influence range.19

The developed ANN model for the single-lateral data set consisted of an input layer (layer 1), two hidden layers (layer 2 and layer 3) with 10 neurons each, and an output layer (layer 4). For the multilateral data set, the developed ANN model consisted of an input layer (layer 1), one hidden layer (layer 2) with 15 neurons, and an output layer (layer 3). For the combined data set, the developed ANN model consisted of an input layer (layer 1), two hidden layers (layer 2 and layer 3) with 10 neurons each, and an output layer (layer 4). Table 11 summarizes the ANN model structure of the single-lateral, multilateral, and combined data sets. It consists of the number of inputs, the number of hidden layers, the number of neurons within each hidden layer, and the number of outputs. This specific model structure was achieved after hundreds of trial-and-error runs to optimize the model outcomes.

Table 11. Final ANN Models’ Structure.

model single-lateral multilateral combined data set
number of inputs 8 7 6
number of neurons in hidden layer 1 10 15 10
number of neurons in hidden layer 2 10 0 10
number of outputs 1 1 1

4. Results and Discussion

Table 12 summarizes the error analysis for both models (FL and ANN) with all three data sets. For error analysis, CC, RMSE, and AAPE were used to evaluate each model. The FL method slightly outperformed the ANN technique as supported by CC values. It showed the highest CC value of 0.95 for single-lateral and combined data sets. However, ANN still showed CC values as good as FL with lower values of RMSE. Nonetheless, this was not the case for the developed models of the multilateral wells, as they showed much worse results compared with the other models. The obtained CCs are 0.85 and 0.80 for FL and ANN, respectively. Figures 16 further evaluate the models by illustrating the regression and actual-predicted assessment plots for the training and the testing data sets for all six models. In Figure 1, the training and testing data sets for the single-lateral wells’ predicted PI using FL (model 1) with very high accuracy (CC equal to 0.94) for both sets of data are given. Similar accuracy is observed when using ANN in model 2 (Figure 2). In Figure 3, using only the wells with multilateral (model 3), FL predicted PI but with a lower accuracy (CC equal to 0.85), and a lower prediction is observed when using ANN in model 4 (CC equal to 0.8) as shown in Figure 4. When combining the single-lateral and multilateral wells in models 5 and 6, the prediction has high accuracy (CC between 0.94 and 0.95) as seen in Figures 5 and 6 using FL and ANN, respectively.

Table 12. Modeling Results Summary.

data set single-lateral multilateral combined data set
model number 1 2 3 4 5 6
method FL ANN FL ANN FL ANN
CC 0.95 0.95 0.85 0.80 0.95 0.94
AAPE 57.7 44.4 25.7 24.6 59.6 65.5
RMSE 21.1 22.4 0.91 0.98 20.0 18.8

Figure 1.

Figure 1

Actual vs predicted regression plot for training and testing data sets (model 1).

Figure 6.

Figure 6

Actual vs predicted regression plot for training and testing data sets (model 6).

Figure 2.

Figure 2

Actual vs predicted regression plot for training and testing data sets (model 2).

Figure 3.

Figure 3

Actual vs predicted regression plot for training and testing data sets (model 3).

Figure 4.

Figure 4

Actual vs predicted regression plot for training and testing data sets (model 4).

Figure 5.

Figure 5

Actual vs predicted regression plot for training and testing data sets (model 5).

4.1. Comparison with Well-Known Correlations

Table 13 compares the values of correlation coefficients (CCs) and root mean square error (RMSE) achieved from the six ML models developed as well as three well-known correlations (Borisov, Joshi, and Furui). It can be deduced that the developed ML models outperformed the well-known correlations for all three different problem sets (single-lateral, multilateral, and combined data sets). It is worth mentioning that among the three well-known correlations, Joshi’s correlation performed the best.

Table 13. CC and RMSE Comparison between the Six Models Using Two ML Methods (FL and ANN) and Three Well-Known Correlations (Borisov, Joshi, and Furui).

  Borisov
Joshi
Furui
FL
ANN
data set/error metric CC RMSE CC RMSE CC RMSE CC RMSE CC RMSE
single-lateral 0.75 50.33 0.77 46.93 0.71 70.39 0.95 21.10 0.95 22.40
multilateral 0.04 7.91 0.05 3.45 0.04 11.56 0.85 0.91 0.8 0.98
Combined 0.26 232.11 0.52 179.50 0.37 74.15 0.95 20.00 0.94 18.80

4.2. Developed Correlation

For models developed using ANN, weight matrices were extracted. Tables 1416 show the weights and biases needed to calculate the PI for each of the three ANN models. ANN neuron calculation can be expressed as in eq 22 to determine n (or wp + bias). Then, n is inserted in the transfer function to calculate a.

4.2. 22

where N is the neuron index, P is the input, and R is the element index in the current element matrix. If another layer is present, this step is repeated with the last calculated a matrix being the new input matrix. The transfer function used to develop the three ANN models was the “Softmax” transfer function. This function is also called the normalized exponential function, and it is expressed as follows

4.2. 23

where z is the input matrix, k is the number of inputs, and i is the input index. Therefore, utilizing eqs 22 and 23, eq 24 is obtained.

4.2. 24

Table 14. ANN Weights and Biases Matrices (ANN for Single-Lateral Wells Data Set).

input layer – hidden layer 1 connection
  input index  
weights 1 2 3 4 5 6 7 8 biases
hidden layer 1 neuron index 1 –3.97 3.64 0.56 0.95 –4.48 –5.00 1.89 1.23 –2.20
2 –1.71 –1.12 –0.95 0.85 1.11 2.10 –2.05 0.19 3.91
3 –0.48 0.00 0.02 –0.02 0.35 0.86 0.30 0.12 –1.25
4 –0.81 –0.20 –1.04 0.87 0.18 0.86 –0.39 0.20 –0.17
5 2.15 –1.23 1.86 –0.21 0.30 1.17 –0.22 1.69 1.25
6 1.80 –0.05 –1.45 1.34 0.10 1.39 1.46 –3.83 2.71
7 0.18 1.45 0.06 –0.07 0.94 1.03 1.16 –0.97 –1.34
8 0.63 1.14 –0.91 –0.19 0.96 –0.02 0.40 –0.80 –1.00
9 1.06 –0.27 0.95 –6.04 –0.42 0.19 –0.53 3.32 –0.59
10 –0.53 –4.15 –0.18 1.12 1.79 –0.29 –1.13 –0.97 –1.06
hidden layer 1 – hidden layer 2 connection
  hidden layer 1 neuron index  
weights 1 2 3 4 5 6 7 8 9 10 biases
hidden layer 2 neuron index 1 –1.66 0.17 0.82 0.45 0.49 1.29 0.00 –0.25 2.26 –0.95 1.30
2 0.32 –0.16 –1.01 0.65 0.46 –0.75 –0.73 –0.80 –0.13 1.21 –0.80
3 –1.47 –0.29 0.94 –0.46 –0.34 –0.76 0.54 1.02 1.07 –1.14 –0.07
4 0.23 –0.05 0.65 –0.07 0.79 –0.33 0.32 –0.97 0.49 0.83 –1.16
5 0.70 –0.91 0.51 0.57 0.17 0.72 –0.75 0.92 0.45 1.07 –1.46
6 0.59 –1.31 –0.53 0.48 0.29 0.13 0.42 –0.43 –1.02 –0.25 –0.11
7 –0.58 –0.23 –0.42 0.25 –0.26 –0.69 0.74 0.39 –0.96 0.33 –0.76
8 0.75 –1.05 –0.02 –0.76 –0.86 0.20 –0.52 –0.31 –0.41 0.71 –0.46
9 –2.28 2.93 –0.60 1.83 –1.18 1.70 –0.14 0.99 0.69 –1.13 2.65
10 3.23 –1.21 –0.75 –0.06 1.85 –1.48 0.21 –0.55 –0.01 1.55 1.45
hidden layer 2 – output layer connection
  hidden layer 2 neuron index  
weights 1 2 3 4 5 6 7 8 9 10 biases
output layer 1 –2.19 1.03 –1.32 0.39 0.63 1.70 0.69 1.55 –1.36 3.86 0.35

Table 16. ANN Weights and Biases Matrices (ANN for Combined Data Set).

Input layer – hidden layer 1 connection
  input index  
weights 1 2 3 4 5 6 biases
hidden layer 1 neuron index 1 –0.97 0.75 –0.01 0.22 0.52 0.57 0.63
2 0.55 0.95 0.24 –0.43 0.39 0.66 –0.72
3 –8.20 7.00 1.67 –9.89 –5.27 1.17 –0.12
4 –1.31 –0.03 –0.56 0.23 –0.46 –0.57 0.82
5 –0.70 1.00 0.37 –0.42 –0.20 0.94 –0.21
6 6.70 2.70 –1.67 –1.58 –8.51 3.35 –3.54
7 –1.38 –1.39 1.30 3.27 1.74 0.46 –1.03
8 2.51 –2.96 –0.76 1.86 1.96 1.16 5.01
9 –0.88 –0.62 0.84 1.01 0.60 0.88 –0.78
10 0.89 –5.65 0.75 3.03 7.31 –7.66 2.09
hidden layer 1 – hidden layer 2 connection
  hidden layer 1 neuron index  
weights 1 2 3 4 5 6 7 8 9 10 biases
hidden layer 2 neuron index 1 0.02 0.42 –0.27 0.90 –0.95 –0.21 0.31 0.47 0.63 –0.23 –1.14
2 –1.13 –0.25 0.57 –0.61 –1.32 –1.46 1.48 –0.04 0.46 –0.16 –0.31
3 0.31 0.39 –3.26 –1.64 0.34 10.12 3.07 –2.45 0.30 –5.97 –1.03
4 0.94 0.60 –0.51 0.31 0.75 0.19 –0.37 0.38 –0.58 –0.70 –1.10
5 0.09 0.66 –0.34 1.36 0.87 –0.26 –1.64 0.91 –0.48 –0.22 –1.51
6 0.94 0.74 7.97 –0.18 0.44 0.80 –1.07 –3.74 –0.52 –0.78 2.38
7 0.61 –0.43 –1.78 2.75 0.97 –1.21 –0.88 –0.79 0.45 4.60 1.97
8 0.76 0.42 0.59 1.21 –0.03 –2.59 –0.75 5.74 –0.40 –0.75 1.55
9 –1.29 –0.74 0.36 –0.05 0.22 –0.64 0.21 –0.77 –0.80 –0.21 –1.35
10 0.04 –0.70 –0.47 –2.13 –0.78 –4.08 –0.32 1.55 0.81 6.10 0.41
hidden layer 2 – output layer connection
  hidden layer 2 neuron index  
weights 1 2 3 4 5 6 7 8 9 10 biases
output layer 1 0.15 0.93 0.32 –0.90 0.44 –1.26 –4.87 4.22 –1.39 2.58 0.28

Table 15. ANN Weights and Biases Matrices (ANN for Multilateral Wells Data Set).

input layer – hidden layer 1 connection
  input index  
weights 1 2 3 4 5 6 7 biases
hidden layer 1 neuron index 1 0.27 0.46 1.04 0.47 –0.55 –0.15 0.45 –0.89
2 –0.53 1.18 –0.17 –0.13 0.98 –0.77 –0.66 –0.96
3 1.34 1.90 0.84 –1.25 –1.20 2.00 2.46 1.83
4 –1.57 0.81 2.57 –0.27 –0.43 0.11 0.61 0.74
5 0.23 –2.08 –1.41 0.56 –0.07 1.90 –0.09 0.91
6 –1.09 –3.19 –2.58 0.93 2.25 –1.30 –4.36 –0.34
7 –2.34 –0.23 –1.27 –2.49 0.81 –0.26 –0.75 –0.85
8 1.40 2.95 1.00 –1.10 –0.16 –3.87 –2.80 –0.35
9 0.01 0.51 0.78 1.81 0.41 –3.38 –2.05 0.27
10 0.85 0.06 1.23 0.70 0.10 –2.05 –2.43 –0.93
11 –0.21 –0.05 –0.91 1.86 0.26 0.14 0.07 –0.09
12 –0.02 –0.97 –0.38 1.09 0.80 1.42 0.13 –0.30
13 0.03 –1.49 1.39 2.85 1.21 1.49 2.82 0.20
14 –2.05 0.27 –0.59 –1.76 1.77 4.75 4.42 0.99
15 0.52 –0.48 –1.81 –0.67 –2.38 –1.38 –2.29 –0.42
hidden layer 1 – output layer connection
  hidden layer 1 neuron index  
weights 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 biases
output layer 1 0.39 0.24 –1.77 1.33 –2.60 0.92 3.65 0.39 –3.07 –1.46 0.24 –1.00 –0.25 0.34 –1.27 –0.56

The developed ML models showed excellent performance compared to other models available in the literature. One major limitation of the well-known correlations, e.g., Borisov and Joshi, is the assumptions made about geometrical factors and flow regimes. This is not the case for the developed ML in this study, as these models require a general set of inputs regardless of the reservoir’s shape and flow. Furthermore, combining the data sets of the single-lateral and multilateral wells to develop one general model that applies to both sets has tackled the low-accuracy issue faced in modeling the productivity index of multilateral wells. Moreover, an attained advantage of the ANN models (models 2, 4, and 6) is the development of a correlation based on the weights, biases, and transfer functions. This correlation can be used to estimate the PI for any new well in the same reservoir using the required set of inputs. Furthermore, the model development process and structure can be implemented to develop correlations for other reservoirs.

5. Conclusions

Multiple correlations are available in the literature to estimate a horizontal well’s productivity. However, each correlation is based on specific assumptions of flow and geometry, which means that they are sometimes not accurate. In such a case, a correlation becomes invalid or gives misleading results. On the basis of the models developed in this study to estimate the productivity index for multilateral oil horizontal wells using machine learning techniques, the following conclusions have been reached:

  • Six ML models were developed to predict the PI for three different data sets (single-lateral, multilateral, and a combination of the two) using two different ML techniques (FL and ANN).

  • The developed ML models showed better performance in predicting PI values compared to other well-known correlations (Borisov, Joshi, and Furui) found in the literature. This is supported by the reported values of CC, RMSE, and AAPE. Also, the utilization of ML models to estimate PI overcomes many of the limitations and assumptions embedded in several widely used correlations.

  • The addition of transformed parameters, such as an input that is a product of two others, has helped in the modeling process providing better results.

  • The integration of alternative correlations (Borisov, Joshi, and Furui) in the development of the ML models as part of the inputs showed improved results. Therefore, such utilization of these correlations, which considers flow regimes and geometrical factors, would add further value and better accuracy to the developed models.

  • Although the multilateral wells PI models showed relatively lower accuracy than the single-lateral wells models, the development of a third model that accounts for both data sets with high accuracy resolved the issue of low accuracy encountered in multilateral wells models.

  • The added value of this study is the developed general and accurate PI estimation model that overcomes the many limitations of several widely used correlations in the industry and can be utilized for single-lateral or multilateral wells.

One of the major limitations of the developed models is they cannot be generalized for any well or field. The developed models are made from a data set of a specific field and are only intended to be utilized for that field where most wells have very similar reservoir characteristics. Nevertheless, the workflow and the model construction and design can be very useful for other fields and can be adopted to obtain accurate prediction models for the data set used to build/train the models.

Acknowledgments

The authors of this article would like to acknowledge the support provided by the King Fahd University of Petroleum & Minerals (KFUPM) to publish this work.

Glossary

Nomenclature

a

semi-major axis of the ellipsoidal drainage area [ft]

AAPE

average absolute percentage error

ANN

artificial neural network

b

reservoir’s length in the well direction [ft]

Bo

oil formation volume factor [bbl/STB]

BPNN

back-propagation neural network

CC

correlation coefficient

CH

geometric factor

FL

fuzzy logic

h

reservoir’s thickness [ft]

Iani

anisotropy factor

k

average permeability [mD]

ke

effective permeability [mD]

kh

horizontal permeability [mD]

kv

vertical permeability [mD]

Le

horizontal well length [ft]

ML

machine learning

μo

oil viscosity [cp]

PI

productivity index [bbl/day/psi]

Pr

reservoir’s pressure [psi]

PSS

pseudosteady state

Pwf

bottom hole pressure [psi]

q

flow rate [bbl/day]

re

reservoir radius [ft]

reh

drainage radius of the horizontal well [ft]

RMSE

root mean square error

rw

well radius [ft]

S

skin damage

w

weight connecting between layers

yb

drainage length perpendicular to the well [ft]

Data Availability Statement

The data sets used and/or analyzed during the current study are available from the corresponding author on reasonable request.

The authors declare no competing financial interest.

Notes

The research does not require any ethical clearance issues.

Due to a production error, the version of this paper that was published ASAP February 11, 2023, contained an error in the second sentence of the abstract. The error was corrected, and the paper was reposted February 21, 2023.

References

  1. Ahmed T.Reservoir Engineering Handbook, 4th ed.; Gulf Professional Publishing: Amsterdam, 2010; pp 457–542. [Google Scholar]
  2. Joshi S. D.Horizontal Well Technology; PennWell Corporation: Tulsa, OK, 1991. [Google Scholar]
  3. Borisov J. P.Oil Production using Horizontal and Multiple Deviation Wells; Nedra: Moscow, 1964. [Google Scholar]
  4. Giger F. M.; Reiss L. H.; Jourdan A. P. In The Reservoir Engineering Aspects of Horizontal Drilling, SPE Annual Technical Conference and Exhibition, 1984.
  5. Joshi S. D. Augmentation of Well Productivity with Slant and Horizontal Wells. J. Pet. Technol. 1988, 40, 729–739. 10.2118/15375-PA. [DOI] [Google Scholar]
  6. Babu D. K.; Odeh A. S. Productivity of a Horizontal Well. SPE Reservoir Eng. 1989, 4, 417–421. 10.2118/18298-PA. [DOI] [Google Scholar]
  7. Economides M. J.; Deimbacher F. X.; Brand C. W.; Heinemann Z. E. Comprehensive Simulation of Horizontal-Well Performance. SPE Form. Eval. 1991, 6, 418–426. 10.2118/20717-PA. [DOI] [Google Scholar]
  8. Renard G.; Dupuy J. M. Formation Damage Effects on Horizontal-Well Flow Efficiency. J. Pet. Technol. 1991, 43, 786–869. 10.2118/19414-PA. [DOI] [Google Scholar]
  9. Butler R. M.Horizontal Wells for the Recovery of Oil, Gas, and Bitumen; Petroleum Society, Canadian Institute of Mining; Metallurgy & Petroleum: Calgary, 1994. [Google Scholar]
  10. Permadi P. In Practical Methods to Forecast Production Performance of Horizontal Wells, SPE Asia Pacific Oil and Gas Conference; OnePetro, 1995.
  11. Elgaghah S. A.; Osisanya S. O.; Tiab D. In A Simple Productivity Equation for Horizontal Wells Based on Drainage Area Concept, SPE Western Regional Meeting; OnePetro, 1996.
  12. Hongen D. In A New Method of Predicting the Productivity and Critical Production Rate Calculation of Horizontal Well, Annual Technical Meeting; OnePetro, 1997.
  13. Lu J.New Productivity Formulae of Horizontal Wells. J. Can. Pet. Technol. 2001, 40 ( (10), ), 10.2118/01-10-03. [DOI]
  14. Furui K.; Zhu D.; Hill A. D. A Rigorous Formation Damage Skin Factor and Reservoir Inflow Model for a Horizontal Well. SPE Prod. Facil. 2003, 18, 151–157. 10.2118/84964-PA. [DOI] [Google Scholar]
  15. Escobar F. H.; Saavedra N. F.; Aranda R. F.; Herrera J. F. In An Improved Correlation to Estimate Productivity Index in Horizontal Wells, SPE Asia Pacific Oil and Gas Conference and Exhibition; OnePetro, 2004.
  16. Farman G. M.; Abdulamir M. R. Formulation of New Equation to Estimate Productivity Index of Horizontal Wells. Iraqi J. Chem. Pet. Eng. 2014, 15, 61–73. 10.31699/IJCPE. [DOI] [Google Scholar]
  17. Alarifi S.; AlNuaim S.; Abdulraheem A. In Productivity Index Prediction for Oil Horizontal Wells Using Different Artificial Intelligence Techniques, SPE Middle East Oil & Gas Show and Conference; OnePetro, 2015.
  18. Al-Mashhad A. S.; Al-Arifi S. A.; Al-Kadem M. S.; Al-Dabbous M. S.; Buhulaigah A. In Multilateral Wells Evaluation Utilizing Artificial Intelligence, SPE Middle East Oil & Gas Show and Conference; OnePetro, 2016.
  19. The MathWorks Inc. . Fuzzy Logic Toolbox TM User’s Guide; The Math Works: Nantick, MA, 2022. [Google Scholar]
  20. Demuth H.; Beale M.. Matlab: Neural Network Toolbox for Use with MATLAB: User’s Guide; The Math Works: Nantick, MA, 2004. [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Data Availability Statement

The data sets used and/or analyzed during the current study are available from the corresponding author on reasonable request.


Articles from ACS Omega are provided here courtesy of American Chemical Society

RESOURCES