Abstract
Horizontal wells are geometrically shaped differently and projected to different flow regimes than vertical wells. Therefore, the existing laws that govern flow and productivity in vertical wells are not applicable to horizontal wells directly. The objective of this paper is to develop machine learning models that predict well productivity index using several reservoir and well inputs. Six models were developed using the actual well rate data from several wells divided into single-lateral wells, multilateral wells, and a combination of single-lateral and multilateral wells. The models are generated using artificial neural networks and fuzzy logic. The inputs used to create the models are the typical inputs used in the correlations and are well-known for any well under production. The results of the established ML models were excellent as suggested by an error analysis performed, reflecting the models to be robust. The error analysis showed high correlation coefficient values (between 0.94 and 0.95) supported by a low estimation error for four models out of six. The added value of this study is the developed general and accurate PI estimation model that overcomes many limitations of several widely used correlations in the industry and can be utilized for single-lateral or multilateral wells.
1. Introduction
Productivity index (PI) is an essential element of the oil and gas well performance evaluation package. It is defined as the quantity of oil produced for a certain drop in pressure. Alternatively, PI is the production flow rate (q) over the drawdown pressure (Δp), which can be expressed as in eq 1. PI unit measurements in field units is bbl/day/psi.
| 1 |
Using Darcy’s definition for single-phase flow in a radial system under steady-state conditions (eq 2), PI for vertical wells can be re-expressed as in eq 3.1
| 2 |
| 3 |
Also, eq 5 expresses the PI for pseudosteady state (PSS) conditions using Darcy’s definition under PSS conditions (eq 4).2
| 4 |
| 5 |
The productivity index is primarily used to evaluate the performance of a certain well, which indeed helps to decide the well’s economic viability. However, this critical parameter is more difficult to obtain for horizontal wells. This is mainly attributed to the different flow behavior observed in horizontal wells, as well as different geometrical factors affecting the well’s productivity.
Generally, horizontal wells perform better than vertical wells because of the larger flow area that horizontal wells can cover. In addition, horizontal wells can be very effective in scenarios where vertical wells would be considered unfavorable, e.g., producing from a thin reservoir. Moreover, an accurate well performance estimate is essential. Therefore, multiple studies attempting to develop PI correlations for horizontal wells are found in the literature. The following table (Table 1) summarizes the different widely used correlations to estimate the PI for horizontal wells and their main assumptions.
Table 1. Summary of Different Widely Used Correlations to Estimate PI for Horizontal Wells.
| correlation | equation | main assumptions | ||
|---|---|---|---|---|
| Borisov3 |
|
isotropic reservoir | ||
| steady-state regime | ||||
| circular drainage area | ||||
| Giger et al.4 |
|
isotropic reservoir | ||
| steady-state regime | ||||
| elliptical geometry | ||||
| Joshi5 |
|
anisotropy | ||
| steady-state regime | ||||
| elliptical geometry | ||||
| Babu and Odeh6 |
|
anisotropy | ||
| pseudosteady state flow | ||||
| box-shaped reservoir | ||||
| damage near wellbore | ||||
| Economides et al.7 |
|
anisotropy | ||
| steady-state flow | ||||
| elliptical drainage area | ||||
| Renard and Dupuy8 |
|
steady-state flow | ||
| circular, elliptical, or rectangular drainage area | ||||
| Butler9 |
|
anisotropy | ||
| box-shaped reservoir | ||||
| Permadi10 |
|
anisotropy | ||
| steady or pseudosteady state | ||||
| elliptical drainage area | ||||
| damage near wellbore | ||||
| Elgaghah et al.11 |
|
splitting up the drainage area into three parts. One semicircle part and two rectangle parts with different lengths | ||
| Hongen12 |
|
steady state | ||
| treated the horizontal well as a vertical well | ||||
| Lu13 |
|
anisotropy | ||
| elliptical shape | ||||
| Furui et al.14 |
|
anisotropy | ||
| radial flow near the wellbore and linear flow far away from the wellbore | ||||
| damage near wellbore | ||||
| Escobar et al.15 |
|
steady state | ||
| based on Joshi’s and Furui’s models | ||||
| Farman and Abdulamir16 |
|
based on Economides’s model | ||
| damage near wellbore |
Alternative approaches to conventional methods were covered in the literature, including utilizing machine learning (ML) techniques. Alarifi et al. introduced several models that employ machine learning methods, e.g., artificial neural network (ANN), fuzzy logic (FL), and functional networks (FN), to model the productivity index of single-lateral horizontal wells. Moreover, Alarifi et al. proceeded to compare the developed ML models with a well-known correlation, e.g., Furui and Joshi, demonstrating an outperformance of ML models over these correlations. Table 2 illustrates a comparison between the error, root mean square error (RMSE), correlation coefficient (CC), average absolute error (AAE), and maximum and minimum absolute errors (AE), obtained from the PI models of several well-known correlations in the industry, and the error obtained from the three developed machine learning models.17
Table 2. Comparison between ML vs Well-Known Correlations17.
| model | Joshi and Economides | Borisov | Giger–Reiss–Jourdan | Renard–Dupuy | Butler | Furui | ANN | FL | FN1 | FN2 |
|---|---|---|---|---|---|---|---|---|---|---|
| RMSE | 43.42 | 156.78 | 63.91 | 89.20 | 66.92 | 66.61 | 22.21 | 18.83 | 23.77 | 23.85 |
| CC | 0.79 | 0.45 | 0.46 | 0.47 | 0.46 | 0.71 | 0.94 | 0.90 | 0.83 | 0.82 |
| AAE | 28.93 | 76.08 | 45.42 | 49.54 | 47.92 | 46.51 | 16.75 | 13.83 | 18.66 | 18.87 |
| maximum AE | 184.02 | 971.90 | 295.13 | 516.53 | 296.87 | 296.5 | 57.21 | 56.63 | 56.16 | 55.88 |
| minimum AE | 0.13 | 0.37 | 0.07 | 0.64 | 0.21 | 0.19 | 0.02 | 0.04 | 0.50 | 0.48 |
Al-Mashhad et al. proposed a machine learning model to evaluate the productivity index of multilateral oil horizontal wells. The ML approach used to develop the model was ANN. As a way to prove the robustness of the developed ML model, Al-Mashhad et al. compared the results with Borisov’s correlation. The results showed superior performance of the ML model over Borisov’s results, as demonstrated in Table 3, which includes average percentage relative error (Er), average absolute percentage relative error (Ea) and its minimum and maximum, spread of the data around the mean (S), and correlation coefficient (R or CC).18
Table 3. Comparison between the ANN Model and Borisov’s Correlation18.
| predicted from ANN |
||||
|---|---|---|---|---|
| measurement | training | test | overall | calculated from Borisov |
| average percentage relative error | –0.48 | –2.93 | –1.17 | 30 |
| average absolute percentage relative error | 5.89 | 12.83 | 7.85 | 50.4 |
| minimum absolute percentage relative error | 0.04 | 0.07 | 0.04 | 3.4 |
| maximum absolute percentage relative error | 38 | 39 | 39 | 143 |
| spread of the data around the mean | 9.4 | 17.1 | 12.1 | 53.8 |
| CC | 0.986 | 0.914 | 0.968 | 0.3 |
Horizontal wells have proved their dominance in the production field over their counterparts, vertical wells. However, horizontal wells are geometrically shaped differently and projected to different flow regimes than what is observed in vertical wells. Therefore, the existing laws that govern the flow and productivity in vertical wells are not applicable to horizontal wells, at least not directly. The literature covers numerous research papers that attempted to find alternative approaches to assess the performance of horizontal wells using the productivity index. Some of these correlations are very well-known, widely accepted, and applied in the industry, such as Joshi’s, Babu, and Odeh and Furui’s correlations. Recently, ML has paved its way through the petroleum industry. This remarkable tool has proven its competence to solve complex modeling problems. It is widely used to predict industry-related parameters, especially those that conventional methods could not solve or solved with high uncertainty, e.g., permeability prediction. This paper will approach the productivity index prediction problem using multiple ML techniques.
The objective of this paper is to employ machine learning techniques to develop a robust model to estimate the PI of horizontal oil wells accurately. The developed model is planned to achieve comparable, if not better, results than the existing well-known correlations, such as Joshi’s and Furui’s. Additionally, some of the developed models used the productivity index estimate calculated using well-known correlations, namely, Borisov, Joshi, and Furui, as inputs in the machine learning (ML) models. Also, to develop a correlation of the resulting weight and bias matrices, which can be then utilized for different data sets of the same input type. The added value of this study is the developed general and accurate PI estimation model that overcomes the many limitations of several widely used correlations in the industry. However, it needs to be taken into account that the ML models are developed for specific ranges of input data sets.
2. Data Preparation
The data under investigation is of two main categories: production data of single-lateral wells and production data of multilateral wells. The data collected for this work is from a valid rate test of several producing horizontal oil wells from a reservoir in the Middle East. All of the wells are currently producing above bubble point pressure. The data consists of 136 different laterals from 28 multilateral wells and 110 rate test data from 110 single-lateral wells (Table 4). Models were developed for each category, including an additional category that includes the combined data of both.
Table 4. Number of Data Points for Each Model.
| data set | single-lateral | multilateral | combined |
|---|---|---|---|
| number of data points | 110 | 136 | 246 |
For single-lateral production data, the available parameters from the 110 wells are reservoir thickness (h), average permeability (k), effective horizontal length (Le), pressure difference between reservoir, wellbore (ΔP), skin factor (s), and oil productivity index (PI). Additional inputs were added as combinations of the provided parameters. For example, PI values using Joshi’s, Furui’s, and Borisov’s correlations were obtained using other inputs, as well as some reservoir fluid, rock, and geometry constants, e.g., viscosity (μo), oil formation volume factor (Bo), anisotropy constant (β), and wellbore radius (w). Additional transformations of the existing inputs were tested as well, such as log(k), ΔP·k, ΔP·log(k), and ΔP/Le.
Statistical analysis was prepared for the inputs to examine each input and investigate if a relationship exists with the output (PI) using the correlation coefficient (CC). Also, another objective was to remove any outlier data that could distort the model development process. A statistical summary of the data is presented in Tables 5 and 6, where Table 5 shows the raw data and Table 6 shows the inputs of transformation and combination.
Table 5. Statistical Summary of Single-Lateral Well Data.
| parameter | h (ft) | k (mD) | Le (ft) | ΔP (psi) | skin | PI (bbl/day/psi) |
|---|---|---|---|---|---|---|
| minimum | 5 | 2 | 165 | 2 | –7.4 | 1.9 |
| maximum | 256 | 5626 | 3600 | 1043 | 180.0 | 385.7 |
| mean | 146 | 632 | 1175 | 193 | 12.6 | 65.0 |
| standard deviation | 53 | 741 | 746 | 220 | 30.8 | 66.1 |
| correlation coefficient with PI | 0.15 | 0.32 | 0.39 | –0.52 | –0.26 | 1.00 |
Table 6. Statistical Analysis of Single-Lateral Well Transformed or Combined Data.
| parameter | log(k) (mD) | ΔP·k (psi mD) | ΔP·log(k) (psi mD) | ΔP/Le (psi/ft) | Borisov (bbl/day/psi) | Joshi (bbl/day/psi) | Furui (bbl/day/psi) |
|---|---|---|---|---|---|---|---|
| minimum | 0.30 | 6.55 × 102 | 6.8 | 0.001 | 0.2 | 0.1 | 0.1 |
| maximum | 3.75 | 1.21 × 106 | 2705.5 | 2.91 | 225.9 | 357.1 | 122.9 |
| mean | 2.60 | 9.54 × 104 | 475.4 | 0.29 | 40.8 | 49.3 | 17.6 |
| standard deviation | 0.50 | 1.77 × 105 | 545.2 | 0.46 | 41.3 | 63.9 | 23.0 |
| correlation coefficient with PI | 0.44 | –0.22 | –0.48 | –0.43 | 0.75 | 0.77 | 0.71 |
Production data of the 28 multilateral wells with 136 different laterals include the average permeability (k), effective horizontal length (Le), wellbore pressure (Pwf), reservoir pressure (Pr), pressure difference between reservoir and wellbore (ΔP), percentage opened of the choke (%choke), and oil productivity index (PI). Additional transformations were added, such as log(k) and ΔP/Le. Table 7 summarizes the statistical analysis of the multilateral well production data.
Table 7. Statistical Summary of Multilateral Well Data.
| parameter | k (mD) | log(k) (mD) | Le (ft) | ΔP (psi) | %choke (%) | Pwf (psi) | Pr (psi) | ΔP/Le (psi/ft) | PI (bbl/day/psi) |
|---|---|---|---|---|---|---|---|---|---|
| minimum | 3 | 0.48 | 2794 | 662 | 5 | 227 | 1831 | 0.09 | 0.7 |
| maximum | 253 | 2.40 | 12 335 | 1948 | 100 | 1621 | 2754 | 0.56 | 8.2 |
| mean | 22 | 1.22 | 6951 | 1325 | 36 | 919 | 2244 | 0.22 | 3.6 |
| standard deviation | 26 | 0.31 | 2327 | 251 | 19 | 242 | 209 | 0.10 | 1.6 |
| correlation coefficient with PI | 0.06 | 0.22 | 0.09 | –0.01 | 0.44 | 0.001 | –0.01 | –0.14 | 1.00 |
To generalize the model, the data of the single-lateral and multilateral wells were combined. However, these two categories do not include the same type of data. Therefore, only data available in both data sets are used. These are average permeability (k), effective horizontal length (Le), pressure difference between reservoir and wellbore (ΔP), and oil productivity index (PI). Once again, further transformations are tested including log(k), ΔP·k, ΔP·log(k), and ΔP/Le.
This combined data set is then tested statistically for the input–output relationship, which would be useful in the model development stage. Table 8 sums up important statistical parameters of the data, such as minimum, maximum, average, standard deviation, and correlation coefficient between input and output.
Table 8. Statistical Summary of the Combined Data.
| parameter | k (mD) | log(k) (mD) | Le (ft) | ΔP (psi) | ΔP·k (psi mD) | ΔP·log(k) (psi mD) | ΔP/Le (psi/ft) | PI (bbl/day/psi) |
|---|---|---|---|---|---|---|---|---|
| minimum | 2 | 0.30 | 165 | 2 | 6.55 × 102 | 7 | 0.001 | 0.7 |
| maximum | 5626 | 3.75 | 12 335 | 1948 | 1.21 × 106 | 3148 | 5.95 | 385.7 |
| mean | 295 | 1.84 | 4369 | 819 | 5.91 × 104 | 1111 | 0.93 | 31.1 |
| standard deviation | 580 | 0.79 | 3393 | 612 | 1.25 × 105 | 780 | 0.96 | 53.7 |
| correlation coefficient with PI | 0.52 | 0.64 | –0.44 | –0.63 | –0.02 | –0.60 | –0.49 | 1.00 |
3. Model Development
Two machine learning techniques (fuzzy logic and artificial neural networks) are used to estimate the productivity index of hundreds of wells. Fuzzy logic applications have significantly increased over the past recent years. This can be attributed to the various applications where this powerful tool can be utilized. This logical system, fuzzy logic, can be defined as “an extension of multivalued logic”. However, in a wider sense, fuzzy logic (FL) is almost synonymous with the theory of fuzzy sets, a theory related to classes of objects with unsharp boundaries in which membership is a matter of degree. This notability, among other AI techniques, comes from many sources, such as its flexibility and easy-to-understand concept.19 FL is known to help deal with uncertainty in engineering. It is an efficient technique to solve complex problems that have high uncertainty and fuzziness of information. It will allow us to handle imprecise knowledge and provide a powerful framework for reasoning.
The FL method selected in this paper for modeling is subtractive clustering. The model is built using an adaptive neurofuzzy inference system, which is a hybrid learning algorithm to identify the membership function parameters of single output, Sugeno-type fuzzy inference systems (FIS). A combination of least-squares and back-propagation gradient descent methods are used for training FIS membership function parameters to model the given set of input/output data. This fast method is best used when there is no clear idea about the data set cluster number. It approximates the data set cluster number and cluster centers, in which later these estimated elements are used to optimize the clustering and model identification process. Ultimately, subtractive clustering generates a fuzzy inference system that describes the input–output behavior.19
Inspired by the biological nervous system, an artificial neural network (ANN) is a system consisting of simple elements (neurons). The network can be trained using weights as a connecting factor between elements to perform a certain function. Based on the connection between elements, different network algorithms can be selected. This robust AI tool has proven its strength in many applications, such as classification, pattern recognition, and identification.
Basically, ANN consists of an input layer, an output layer, and hidden or more layers. The input layer is connected to a hidden layer, and the hidden layer is connected to the next hidden layer if available, or if there is only one hidden layer, it is connected to the output layer. The connection between the layers is achieved through weights and biases. In a typical neural network, P refers to the input, R is the number of inputs, w is the weight connecting between layers (e.g., input-to-neuron connection), b is the neuron bias, and “n” is assigned as “Wp + b” and inserted in “f”, which represents a transfer function. The objective of a transfer function is to take “n” and produce an output “a”, which will be assigned differently to each available neuron. The produced values of “a” will be ultimately used to compute the output.20
This work employs the back-propagation neural network (BPNN) technique in the modeling process. In simple terms, BPNN is based on generalizing the Widrow–Hoff learning rules. It follows a gradient descent algorithm in which it moves the weights along the negative of the gradient performance function to achieve a particular solution. The network architecture utilized was feedforward. It is a common network architecture used with a back-propagation algorithm.20
The process of model development consisted of six models, in which two models were developed for each data set (single-lateral, multilateral, and combined). Table 9 sums up the ML methods used for each model. These models were developed using the aforementioned ML techniques, FL and ANN. For each model, the data was split into 70% for training and 30% for testing for fuzzy logic models and 70% for training, 15% for internal validation, and 15% for testing for ANN models.
Table 9. Summary of the ML Methods Used for Each Model.
| data set | single-lateral | multilateral | combined data set | |||
|---|---|---|---|---|---|---|
| model number | 1 | 2 | 3 | 4 | 5 | 6 |
| ML method | FL | ANN | FL | ANN | FL | ANN |
To assess a model’s quality, several parameters can be utilized. A correlation coefficient (CC) is a statistical parameter that tests the linear relationship between actual values and predicted values. Statistically, CC values can be expressed in the range of −1 to 1, with values closer to 1 indicating a strong positive relationship, while values approaching −1 represent a negative relation. In terms of the use of CC in model assessment, the closer the value of CC is to 1, the lower the error between actual and predicted values. Another parameter used to judge the goodness of a model is root mean square error (RMSE). RMSE is the average squared difference between the actual values and their corresponding predicted values and can be described mathematically as follows (eq 20)
| 20 |
where n is the number of points in the data set. Lower values of RMSE are desired for better models. An additional parameter that was employed to evaluate the model is the average absolute percentage error (AAPE). AAPE is the average of the percentage difference between actual and predicted values divided by their actual values. To have a better model, the goal is to aim for lower values of AAPE. AAPE can be expressed as (eq 21)
| 21 |
Hundreds of trials were run, including changing the model inputs used, until an optimum model is found for each data set (lowest prediction error). The final inputs that were used to give the best results for each model are listed in Table 10.
Table 10. Final List of Inputs Used for Each Model.
| model | inputs | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| 1 | h | S | Le | ΔP | log k | ΔP·log(k) | ΔP/Le | Joshi | Furui | Borisov | |
| 2 | h | S | Le | ΔP | log k | ΔP·log(k) | ΔP/Le | Joshi | |||
| 3 | k | Le | ΔP | log k | Pwf | Pr | %choke | ||||
| 4 | k | Le | ΔP | ΔP/Le | Pwf | Pr | %choke | ||||
| 5 | Le | ΔP | log k | ΔP·log(k) | ΔP/Le | ||||||
| 6 | Le | ΔP | log k | ΔP·k | ΔP·log(k) | ΔP/Le | |||||
The use of transformed parameters, such as an input that is a product of two others, as model inputs helped in the modeling process, providing better results. Also, using alternative correlations (Borisov, Joshi, and Furui) in the development of the ML models as part of the inputs showed improved results. Such utilization of these correlations, which considers flow regimes and geometrical factors, would add further value and better accuracy to the developed models.
The fuzzy logic systems used to develop the model were Sugeno fuzzy systems, specifically, using a subtractive clustering algorithm. The process of this method starts with the calculation of the likelihood for each data point to define a cluster center according to the surrounding data point density. Second, the data point with the highest potential to be the first cluster center is chosen and the remaining data points surrounding it are removed. Then, the next data point with the highest potential from among the remaining data points is chosen and these steps are repeated until all of the data is within a cluster center influence range.19
The developed ANN model for the single-lateral data set consisted of an input layer (layer 1), two hidden layers (layer 2 and layer 3) with 10 neurons each, and an output layer (layer 4). For the multilateral data set, the developed ANN model consisted of an input layer (layer 1), one hidden layer (layer 2) with 15 neurons, and an output layer (layer 3). For the combined data set, the developed ANN model consisted of an input layer (layer 1), two hidden layers (layer 2 and layer 3) with 10 neurons each, and an output layer (layer 4). Table 11 summarizes the ANN model structure of the single-lateral, multilateral, and combined data sets. It consists of the number of inputs, the number of hidden layers, the number of neurons within each hidden layer, and the number of outputs. This specific model structure was achieved after hundreds of trial-and-error runs to optimize the model outcomes.
Table 11. Final ANN Models’ Structure.
| model | single-lateral | multilateral | combined data set |
|---|---|---|---|
| number of inputs | 8 | 7 | 6 |
| number of neurons in hidden layer 1 | 10 | 15 | 10 |
| number of neurons in hidden layer 2 | 10 | 0 | 10 |
| number of outputs | 1 | 1 | 1 |
4. Results and Discussion
Table 12 summarizes the error analysis for both models (FL and ANN) with all three data sets. For error analysis, CC, RMSE, and AAPE were used to evaluate each model. The FL method slightly outperformed the ANN technique as supported by CC values. It showed the highest CC value of 0.95 for single-lateral and combined data sets. However, ANN still showed CC values as good as FL with lower values of RMSE. Nonetheless, this was not the case for the developed models of the multilateral wells, as they showed much worse results compared with the other models. The obtained CCs are 0.85 and 0.80 for FL and ANN, respectively. Figures 1–6 further evaluate the models by illustrating the regression and actual-predicted assessment plots for the training and the testing data sets for all six models. In Figure 1, the training and testing data sets for the single-lateral wells’ predicted PI using FL (model 1) with very high accuracy (CC equal to 0.94) for both sets of data are given. Similar accuracy is observed when using ANN in model 2 (Figure 2). In Figure 3, using only the wells with multilateral (model 3), FL predicted PI but with a lower accuracy (CC equal to 0.85), and a lower prediction is observed when using ANN in model 4 (CC equal to 0.8) as shown in Figure 4. When combining the single-lateral and multilateral wells in models 5 and 6, the prediction has high accuracy (CC between 0.94 and 0.95) as seen in Figures 5 and 6 using FL and ANN, respectively.
Table 12. Modeling Results Summary.
| data set | single-lateral | multilateral | combined data set | |||
|---|---|---|---|---|---|---|
| model number | 1 | 2 | 3 | 4 | 5 | 6 |
| method | FL | ANN | FL | ANN | FL | ANN |
| CC | 0.95 | 0.95 | 0.85 | 0.80 | 0.95 | 0.94 |
| AAPE | 57.7 | 44.4 | 25.7 | 24.6 | 59.6 | 65.5 |
| RMSE | 21.1 | 22.4 | 0.91 | 0.98 | 20.0 | 18.8 |
Figure 1.
Actual vs predicted regression plot for training and testing data sets (model 1).
Figure 6.
Actual vs predicted regression plot for training and testing data sets (model 6).
Figure 2.
Actual vs predicted regression plot for training and testing data sets (model 2).
Figure 3.
Actual vs predicted regression plot for training and testing data sets (model 3).
Figure 4.
Actual vs predicted regression plot for training and testing data sets (model 4).
Figure 5.
Actual vs predicted regression plot for training and testing data sets (model 5).
4.1. Comparison with Well-Known Correlations
Table 13 compares the values of correlation coefficients (CCs) and root mean square error (RMSE) achieved from the six ML models developed as well as three well-known correlations (Borisov, Joshi, and Furui). It can be deduced that the developed ML models outperformed the well-known correlations for all three different problem sets (single-lateral, multilateral, and combined data sets). It is worth mentioning that among the three well-known correlations, Joshi’s correlation performed the best.
Table 13. CC and RMSE Comparison between the Six Models Using Two ML Methods (FL and ANN) and Three Well-Known Correlations (Borisov, Joshi, and Furui).
| Borisov |
Joshi |
Furui |
FL |
ANN |
||||||
|---|---|---|---|---|---|---|---|---|---|---|
| data set/error metric | CC | RMSE | CC | RMSE | CC | RMSE | CC | RMSE | CC | RMSE |
| single-lateral | 0.75 | 50.33 | 0.77 | 46.93 | 0.71 | 70.39 | 0.95 | 21.10 | 0.95 | 22.40 |
| multilateral | 0.04 | 7.91 | 0.05 | 3.45 | 0.04 | 11.56 | 0.85 | 0.91 | 0.8 | 0.98 |
| Combined | 0.26 | 232.11 | 0.52 | 179.50 | 0.37 | 74.15 | 0.95 | 20.00 | 0.94 | 18.80 |
4.2. Developed Correlation
For models developed using ANN, weight matrices were extracted. Tables 14–16 show the weights and biases needed to calculate the PI for each of the three ANN models. ANN neuron calculation can be expressed as in eq 22 to determine n (or wp + bias). Then, n is inserted in the transfer function to calculate a.
| 22 |
where N is the neuron index, P is the input, and R is the element index in the current element matrix. If another layer is present, this step is repeated with the last calculated a matrix being the new input matrix. The transfer function used to develop the three ANN models was the “Softmax” transfer function. This function is also called the normalized exponential function, and it is expressed as follows
| 23 |
where z is the input matrix, k is the number of inputs, and i is the input index. Therefore, utilizing eqs 22 and 23, eq 24 is obtained.
| 24 |
Table 14. ANN Weights and Biases Matrices (ANN for Single-Lateral Wells Data Set).
| input
layer – hidden layer 1 connection | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| input index | ||||||||||
| weights | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | biases | |
| hidden layer 1 neuron index | 1 | –3.97 | 3.64 | 0.56 | 0.95 | –4.48 | –5.00 | 1.89 | 1.23 | –2.20 |
| 2 | –1.71 | –1.12 | –0.95 | 0.85 | 1.11 | 2.10 | –2.05 | 0.19 | 3.91 | |
| 3 | –0.48 | 0.00 | 0.02 | –0.02 | 0.35 | 0.86 | 0.30 | 0.12 | –1.25 | |
| 4 | –0.81 | –0.20 | –1.04 | 0.87 | 0.18 | 0.86 | –0.39 | 0.20 | –0.17 | |
| 5 | 2.15 | –1.23 | 1.86 | –0.21 | 0.30 | 1.17 | –0.22 | 1.69 | 1.25 | |
| 6 | 1.80 | –0.05 | –1.45 | 1.34 | 0.10 | 1.39 | 1.46 | –3.83 | 2.71 | |
| 7 | 0.18 | 1.45 | 0.06 | –0.07 | 0.94 | 1.03 | 1.16 | –0.97 | –1.34 | |
| 8 | 0.63 | 1.14 | –0.91 | –0.19 | 0.96 | –0.02 | 0.40 | –0.80 | –1.00 | |
| 9 | 1.06 | –0.27 | 0.95 | –6.04 | –0.42 | 0.19 | –0.53 | 3.32 | –0.59 | |
| 10 | –0.53 | –4.15 | –0.18 | 1.12 | 1.79 | –0.29 | –1.13 | –0.97 | –1.06 | |
| hidden layer 1 –
hidden layer 2 connection | ||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| hidden layer 1 neuron index | ||||||||||||
| weights | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 | 10 | biases | |
| hidden layer 2 neuron index | 1 | –1.66 | 0.17 | 0.82 | 0.45 | 0.49 | 1.29 | 0.00 | –0.25 | 2.26 | –0.95 | 1.30 |
| 2 | 0.32 | –0.16 | –1.01 | 0.65 | 0.46 | –0.75 | –0.73 | –0.80 | –0.13 | 1.21 | –0.80 | |
| 3 | –1.47 | –0.29 | 0.94 | –0.46 | –0.34 | –0.76 | 0.54 | 1.02 | 1.07 | –1.14 | –0.07 | |
| 4 | 0.23 | –0.05 | 0.65 | –0.07 | 0.79 | –0.33 | 0.32 | –0.97 | 0.49 | 0.83 | –1.16 | |
| 5 | 0.70 | –0.91 | 0.51 | 0.57 | 0.17 | 0.72 | –0.75 | 0.92 | 0.45 | 1.07 | –1.46 | |
| 6 | 0.59 | –1.31 | –0.53 | 0.48 | 0.29 | 0.13 | 0.42 | –0.43 | –1.02 | –0.25 | –0.11 | |
| 7 | –0.58 | –0.23 | –0.42 | 0.25 | –0.26 | –0.69 | 0.74 | 0.39 | –0.96 | 0.33 | –0.76 | |
| 8 | 0.75 | –1.05 | –0.02 | –0.76 | –0.86 | 0.20 | –0.52 | –0.31 | –0.41 | 0.71 | –0.46 | |
| 9 | –2.28 | 2.93 | –0.60 | 1.83 | –1.18 | 1.70 | –0.14 | 0.99 | 0.69 | –1.13 | 2.65 | |
| 10 | 3.23 | –1.21 | –0.75 | –0.06 | 1.85 | –1.48 | 0.21 | –0.55 | –0.01 | 1.55 | 1.45 | |
| hidden layer 2 – output layer connection | ||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| hidden layer 2 neuron index | ||||||||||||
| weights | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 | 10 | biases | |
| output layer | 1 | –2.19 | 1.03 | –1.32 | 0.39 | 0.63 | 1.70 | 0.69 | 1.55 | –1.36 | 3.86 | 0.35 |
Table 16. ANN Weights and Biases Matrices (ANN for Combined Data Set).
| Input layer – hidden layer 1 connection | ||||||||
|---|---|---|---|---|---|---|---|---|
| input index | ||||||||
| weights | 1 | 2 | 3 | 4 | 5 | 6 | biases | |
| hidden layer 1 neuron index | 1 | –0.97 | 0.75 | –0.01 | 0.22 | 0.52 | 0.57 | 0.63 |
| 2 | 0.55 | 0.95 | 0.24 | –0.43 | 0.39 | 0.66 | –0.72 | |
| 3 | –8.20 | 7.00 | 1.67 | –9.89 | –5.27 | 1.17 | –0.12 | |
| 4 | –1.31 | –0.03 | –0.56 | 0.23 | –0.46 | –0.57 | 0.82 | |
| 5 | –0.70 | 1.00 | 0.37 | –0.42 | –0.20 | 0.94 | –0.21 | |
| 6 | 6.70 | 2.70 | –1.67 | –1.58 | –8.51 | 3.35 | –3.54 | |
| 7 | –1.38 | –1.39 | 1.30 | 3.27 | 1.74 | 0.46 | –1.03 | |
| 8 | 2.51 | –2.96 | –0.76 | 1.86 | 1.96 | 1.16 | 5.01 | |
| 9 | –0.88 | –0.62 | 0.84 | 1.01 | 0.60 | 0.88 | –0.78 | |
| 10 | 0.89 | –5.65 | 0.75 | 3.03 | 7.31 | –7.66 | 2.09 | |
| hidden layer 1 – hidden layer 2 connection | ||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| hidden layer 1 neuron index | ||||||||||||
| weights | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 | 10 | biases | |
| hidden layer 2 neuron index | 1 | 0.02 | 0.42 | –0.27 | 0.90 | –0.95 | –0.21 | 0.31 | 0.47 | 0.63 | –0.23 | –1.14 |
| 2 | –1.13 | –0.25 | 0.57 | –0.61 | –1.32 | –1.46 | 1.48 | –0.04 | 0.46 | –0.16 | –0.31 | |
| 3 | 0.31 | 0.39 | –3.26 | –1.64 | 0.34 | 10.12 | 3.07 | –2.45 | 0.30 | –5.97 | –1.03 | |
| 4 | 0.94 | 0.60 | –0.51 | 0.31 | 0.75 | 0.19 | –0.37 | 0.38 | –0.58 | –0.70 | –1.10 | |
| 5 | 0.09 | 0.66 | –0.34 | 1.36 | 0.87 | –0.26 | –1.64 | 0.91 | –0.48 | –0.22 | –1.51 | |
| 6 | 0.94 | 0.74 | 7.97 | –0.18 | 0.44 | 0.80 | –1.07 | –3.74 | –0.52 | –0.78 | 2.38 | |
| 7 | 0.61 | –0.43 | –1.78 | 2.75 | 0.97 | –1.21 | –0.88 | –0.79 | 0.45 | 4.60 | 1.97 | |
| 8 | 0.76 | 0.42 | 0.59 | 1.21 | –0.03 | –2.59 | –0.75 | 5.74 | –0.40 | –0.75 | 1.55 | |
| 9 | –1.29 | –0.74 | 0.36 | –0.05 | 0.22 | –0.64 | 0.21 | –0.77 | –0.80 | –0.21 | –1.35 | |
| 10 | 0.04 | –0.70 | –0.47 | –2.13 | –0.78 | –4.08 | –0.32 | 1.55 | 0.81 | 6.10 | 0.41 | |
| hidden layer 2 –
output layer connection | ||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| hidden layer 2 neuron index | ||||||||||||
| weights | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 | 10 | biases | |
| output layer | 1 | 0.15 | 0.93 | 0.32 | –0.90 | 0.44 | –1.26 | –4.87 | 4.22 | –1.39 | 2.58 | 0.28 |
Table 15. ANN Weights and Biases Matrices (ANN for Multilateral Wells Data Set).
| input
layer – hidden layer 1 connection | |||||||||
|---|---|---|---|---|---|---|---|---|---|
| input index | |||||||||
| weights | 1 | 2 | 3 | 4 | 5 | 6 | 7 | biases | |
| hidden layer 1 neuron index | 1 | 0.27 | 0.46 | 1.04 | 0.47 | –0.55 | –0.15 | 0.45 | –0.89 |
| 2 | –0.53 | 1.18 | –0.17 | –0.13 | 0.98 | –0.77 | –0.66 | –0.96 | |
| 3 | 1.34 | 1.90 | 0.84 | –1.25 | –1.20 | 2.00 | 2.46 | 1.83 | |
| 4 | –1.57 | 0.81 | 2.57 | –0.27 | –0.43 | 0.11 | 0.61 | 0.74 | |
| 5 | 0.23 | –2.08 | –1.41 | 0.56 | –0.07 | 1.90 | –0.09 | 0.91 | |
| 6 | –1.09 | –3.19 | –2.58 | 0.93 | 2.25 | –1.30 | –4.36 | –0.34 | |
| 7 | –2.34 | –0.23 | –1.27 | –2.49 | 0.81 | –0.26 | –0.75 | –0.85 | |
| 8 | 1.40 | 2.95 | 1.00 | –1.10 | –0.16 | –3.87 | –2.80 | –0.35 | |
| 9 | 0.01 | 0.51 | 0.78 | 1.81 | 0.41 | –3.38 | –2.05 | 0.27 | |
| 10 | 0.85 | 0.06 | 1.23 | 0.70 | 0.10 | –2.05 | –2.43 | –0.93 | |
| 11 | –0.21 | –0.05 | –0.91 | 1.86 | 0.26 | 0.14 | 0.07 | –0.09 | |
| 12 | –0.02 | –0.97 | –0.38 | 1.09 | 0.80 | 1.42 | 0.13 | –0.30 | |
| 13 | 0.03 | –1.49 | 1.39 | 2.85 | 1.21 | 1.49 | 2.82 | 0.20 | |
| 14 | –2.05 | 0.27 | –0.59 | –1.76 | 1.77 | 4.75 | 4.42 | 0.99 | |
| 15 | 0.52 | –0.48 | –1.81 | –0.67 | –2.38 | –1.38 | –2.29 | –0.42 | |
| hidden layer 1 –
output layer connection | |||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| hidden layer 1 neuron index | |||||||||||||||||
| weights | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 | 10 | 11 | 12 | 13 | 14 | 15 | biases | |
| output layer | 1 | 0.39 | 0.24 | –1.77 | 1.33 | –2.60 | 0.92 | 3.65 | 0.39 | –3.07 | –1.46 | 0.24 | –1.00 | –0.25 | 0.34 | –1.27 | –0.56 |
The developed ML models showed excellent performance compared to other models available in the literature. One major limitation of the well-known correlations, e.g., Borisov and Joshi, is the assumptions made about geometrical factors and flow regimes. This is not the case for the developed ML in this study, as these models require a general set of inputs regardless of the reservoir’s shape and flow. Furthermore, combining the data sets of the single-lateral and multilateral wells to develop one general model that applies to both sets has tackled the low-accuracy issue faced in modeling the productivity index of multilateral wells. Moreover, an attained advantage of the ANN models (models 2, 4, and 6) is the development of a correlation based on the weights, biases, and transfer functions. This correlation can be used to estimate the PI for any new well in the same reservoir using the required set of inputs. Furthermore, the model development process and structure can be implemented to develop correlations for other reservoirs.
5. Conclusions
Multiple correlations are available in the literature to estimate a horizontal well’s productivity. However, each correlation is based on specific assumptions of flow and geometry, which means that they are sometimes not accurate. In such a case, a correlation becomes invalid or gives misleading results. On the basis of the models developed in this study to estimate the productivity index for multilateral oil horizontal wells using machine learning techniques, the following conclusions have been reached:
Six ML models were developed to predict the PI for three different data sets (single-lateral, multilateral, and a combination of the two) using two different ML techniques (FL and ANN).
The developed ML models showed better performance in predicting PI values compared to other well-known correlations (Borisov, Joshi, and Furui) found in the literature. This is supported by the reported values of CC, RMSE, and AAPE. Also, the utilization of ML models to estimate PI overcomes many of the limitations and assumptions embedded in several widely used correlations.
The addition of transformed parameters, such as an input that is a product of two others, has helped in the modeling process providing better results.
The integration of alternative correlations (Borisov, Joshi, and Furui) in the development of the ML models as part of the inputs showed improved results. Therefore, such utilization of these correlations, which considers flow regimes and geometrical factors, would add further value and better accuracy to the developed models.
Although the multilateral wells PI models showed relatively lower accuracy than the single-lateral wells models, the development of a third model that accounts for both data sets with high accuracy resolved the issue of low accuracy encountered in multilateral wells models.
The added value of this study is the developed general and accurate PI estimation model that overcomes the many limitations of several widely used correlations in the industry and can be utilized for single-lateral or multilateral wells.
One of the major limitations of the developed models is they cannot be generalized for any well or field. The developed models are made from a data set of a specific field and are only intended to be utilized for that field where most wells have very similar reservoir characteristics. Nevertheless, the workflow and the model construction and design can be very useful for other fields and can be adopted to obtain accurate prediction models for the data set used to build/train the models.
Acknowledgments
The authors of this article would like to acknowledge the support provided by the King Fahd University of Petroleum & Minerals (KFUPM) to publish this work.
Glossary
Nomenclature
- a
semi-major axis of the ellipsoidal drainage area [ft]
- AAPE
average absolute percentage error
- ANN
artificial neural network
- b
reservoir’s length in the well direction [ft]
- Bo
oil formation volume factor [bbl/STB]
- BPNN
back-propagation neural network
- CC
correlation coefficient
- CH
geometric factor
- FL
fuzzy logic
- h
reservoir’s thickness [ft]
- Iani
anisotropy factor
- k
average permeability [mD]
- ke
effective permeability [mD]
- kh
horizontal permeability [mD]
- kv
vertical permeability [mD]
- Le
horizontal well length [ft]
- ML
machine learning
- μo
oil viscosity [cp]
- PI
productivity index [bbl/day/psi]
- Pr
reservoir’s pressure [psi]
- PSS
pseudosteady state
- Pwf
bottom hole pressure [psi]
- q
flow rate [bbl/day]
- re
reservoir radius [ft]
- reh
drainage radius of the horizontal well [ft]
- RMSE
root mean square error
- rw
well radius [ft]
- S
skin damage
- w
weight connecting between layers
- yb
drainage length perpendicular to the well [ft]
Data Availability Statement
The data sets used and/or analyzed during the current study are available from the corresponding author on reasonable request.
The authors declare no competing financial interest.
Notes
The research does not require any ethical clearance issues.
Due to a production error, the version of this paper that was published ASAP February 11, 2023, contained an error in the second sentence of the abstract. The error was corrected, and the paper was reposted February 21, 2023.
References
- Ahmed T.Reservoir Engineering Handbook, 4th ed.; Gulf Professional Publishing: Amsterdam, 2010; pp 457–542. [Google Scholar]
- Joshi S. D.Horizontal Well Technology; PennWell Corporation: Tulsa, OK, 1991. [Google Scholar]
- Borisov J. P.Oil Production using Horizontal and Multiple Deviation Wells; Nedra: Moscow, 1964. [Google Scholar]
- Giger F. M.; Reiss L. H.; Jourdan A. P. In The Reservoir Engineering Aspects of Horizontal Drilling, SPE Annual Technical Conference and Exhibition, 1984.
- Joshi S. D. Augmentation of Well Productivity with Slant and Horizontal Wells. J. Pet. Technol. 1988, 40, 729–739. 10.2118/15375-PA. [DOI] [Google Scholar]
- Babu D. K.; Odeh A. S. Productivity of a Horizontal Well. SPE Reservoir Eng. 1989, 4, 417–421. 10.2118/18298-PA. [DOI] [Google Scholar]
- Economides M. J.; Deimbacher F. X.; Brand C. W.; Heinemann Z. E. Comprehensive Simulation of Horizontal-Well Performance. SPE Form. Eval. 1991, 6, 418–426. 10.2118/20717-PA. [DOI] [Google Scholar]
- Renard G.; Dupuy J. M. Formation Damage Effects on Horizontal-Well Flow Efficiency. J. Pet. Technol. 1991, 43, 786–869. 10.2118/19414-PA. [DOI] [Google Scholar]
- Butler R. M.Horizontal Wells for the Recovery of Oil, Gas, and Bitumen; Petroleum Society, Canadian Institute of Mining; Metallurgy & Petroleum: Calgary, 1994. [Google Scholar]
- Permadi P. In Practical Methods to Forecast Production Performance of Horizontal Wells, SPE Asia Pacific Oil and Gas Conference; OnePetro, 1995.
- Elgaghah S. A.; Osisanya S. O.; Tiab D. In A Simple Productivity Equation for Horizontal Wells Based on Drainage Area Concept, SPE Western Regional Meeting; OnePetro, 1996.
- Hongen D. In A New Method of Predicting the Productivity and Critical Production Rate Calculation of Horizontal Well, Annual Technical Meeting; OnePetro, 1997.
- Lu J.New Productivity Formulae of Horizontal Wells. J. Can. Pet. Technol. 2001, 40 ( (10), ), 10.2118/01-10-03. [DOI]
- Furui K.; Zhu D.; Hill A. D. A Rigorous Formation Damage Skin Factor and Reservoir Inflow Model for a Horizontal Well. SPE Prod. Facil. 2003, 18, 151–157. 10.2118/84964-PA. [DOI] [Google Scholar]
- Escobar F. H.; Saavedra N. F.; Aranda R. F.; Herrera J. F. In An Improved Correlation to Estimate Productivity Index in Horizontal Wells, SPE Asia Pacific Oil and Gas Conference and Exhibition; OnePetro, 2004.
- Farman G. M.; Abdulamir M. R. Formulation of New Equation to Estimate Productivity Index of Horizontal Wells. Iraqi J. Chem. Pet. Eng. 2014, 15, 61–73. 10.31699/IJCPE. [DOI] [Google Scholar]
- Alarifi S.; AlNuaim S.; Abdulraheem A. In Productivity Index Prediction for Oil Horizontal Wells Using Different Artificial Intelligence Techniques, SPE Middle East Oil & Gas Show and Conference; OnePetro, 2015.
- Al-Mashhad A. S.; Al-Arifi S. A.; Al-Kadem M. S.; Al-Dabbous M. S.; Buhulaigah A. In Multilateral Wells Evaluation Utilizing Artificial Intelligence, SPE Middle East Oil & Gas Show and Conference; OnePetro, 2016.
- The MathWorks Inc. . Fuzzy Logic Toolbox TM User’s Guide; The Math Works: Nantick, MA, 2022. [Google Scholar]
- Demuth H.; Beale M.. Matlab: Neural Network Toolbox for Use with MATLAB: User’s Guide; The Math Works: Nantick, MA, 2004. [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Data Availability Statement
The data sets used and/or analyzed during the current study are available from the corresponding author on reasonable request.















