Abstract
Understanding the microstructure–property relationships of porous media is of great practical significance, based on which macroscopic physical properties can be directly derived from measurable microstructural informatics. However, establishing reliable microstructure–property mappings in an explicit manner is difficult, due to the intricacy, stochasticity, and heterogeneity of porous microstructures. In this paper, a data-driven computational framework is presented to investigate the inherent microstructure–permeability linkage for natural porous rocks, where multiple techniques are integrated together, including microscopy imaging, stochastic reconstruction, microstructural characterization, pore-scale simulation, feature selection, and data-driven modeling. A large number of 3D digital rocks with a wide porosity range are acquired from microscopy imaging and stochastic reconstruction techniques. A broad variety of morphological descriptors are used to quantitatively characterize pore microstructures from different perspectives, and they compose the raw feature pool for feature selection. High-fidelity lattice Boltzmann simulations are conducted to resolve fluid flow passing through porous media, from which reliable permeability references are obtained. The optimal feature set that best represents permeability is identified through a performance-oriented feature selection process, upon which a cost-effective surrogate model is rapidly fitted to approximate the microstructure-permeability mapping via data-driven modeling. This surrogate model exhibits great advantages over empirical/analytical formulas in terms of prediction accuracy and generalization capacity, which can predict reliable permeability values spanning four orders of magnitude. Besides, feature selection also greatly enhances the interpretability of the data-driven prediction model, from which new insights into the mechanism of how microstructural characteristics determine intrinsic permeability are obtained.
Keywords: Porous rocks, Permeability prediction, Microstructural characterization, Lattice Boltzmann simulation, Feature selection, Data-driven modeling
Introduction
Permeability quantifies the ability of a porous medium to transmit fluid and serves as a fundamental characteristic for the transport behavior of fluid flow inside porous media [1, 8]. It plays a critical role in such geological applications as oil and gas recovery, geothermal energy exploitation, CO underground storage, radioactive waste disposal, and contaminant hydrogeology. The permeable pore spaces in different geologic materials are often highly distinctive, leading to an extremely broad range of permeability values that vary up to 13 orders of magnitude [73]. The macroscopic physical properties of porous media [10, 35, 41, 47, 74, 75] strongly depend on the microstructural characteristics, such that the hydraulic, mechanical, electrical, and thermal properties of porous media can be evaluated/estimated from the measurable microstructural informatics, at least in principle. Indeed, the microstructure–property relationship is one of the most fundamental challenges in porous media research. However, the intricacy, stochasticity, and heterogeneity inherent in natural porous rocks make it difficult to accurately and rapidly evaluate permeability, especially for tight rocks with low porosity. Therefore, a deep insight into the microstructure–permeability mapping is desirable and has attracted extensive research efforts, to develop a reliable and efficient method for permeability prediction [13, 34, 80, 83, 100].
Laboratory measurement is the routine way to determine permeability, where fluid flow is driven by a constant pressure difference to pass through a rock core and permeability is then evaluated according to Darcy’s law when the fluid flow reaches the steady state [8]. In practical applications, long waiting time and high cost are the main limitations of experimental measurement, especially for tight rocks. In addition to the experimental measurement, analytical and empirical models have also been developed to predict permeability of porous media, such as the well-known Kozeny–Carman relation [21, 35] and many variants derived from it [10, 11, 28]. Generally, these models rely on specific microstructural characteristics of porous media, such as porosity, specific surface area, tortuosity, characteristic length, pore size, constriction factor, and fractal dimension, among others. Despite the simplicity and convenience in practical applications, analytical models are often overly idealized and empirical models usually contain adjustable parameters to accommodate uncertainty. As a result, they are restricted to some specific pore microstructures and for natural rocks with complicated pore networks, their prediction accuracy drops significantly (and can become unacceptable due to excessive errors).
In recent years, the digital rock physics (DRP) technique has progressed rapidly, offering an alternative to laboratory measurement and analytical/empirical models for permeability evaluation. The DRP approach uses advanced microscopy imaging techniques [4, 17], such as X-ray micro-computed tomography (micro-CT) and focused ion beam scanning electron microscopy (FIB-SEM), to obtain 3D geometries of pore microstructures, on which high-fidelity numerical simulations are performed to evaluate various transport properties and investigate specific physical phenomena [3, 16]. The DRP approach is convenient and promising for microstructural characterization and petrophysical property evaluation [3, 34].
Pore-network modeling (PNM) and direct numerical simulation (DNS) are the two primary pore-scale computing approaches to mimicking transport processes occurring inside porous rocks. According to some specific criteria [16, 113], PNM simplifies the complicated pore space into a topologically representative network of pore bodies interconnected by pore throats with ideal shapes (such as sphere and cylinder). The transport behaviors within each network element are described by semi-analytical laws (such as Hagen–Poiseuille law), which greatly reduces the computational cost and enables multi-scale modeling to incorporate strong heterogeneity in large volumes. PNM is widely used for capillary-controlled transport processes. However, due to the simplification of complicated pore space, PNM may produce inaccurate estimations. To date, it remains a major challenge to correctly identify the key microstructural features that are critical for effective PNM estimation and the less important ones that can be safely ignored to reduce computational complexity [113].
By contrast, DNS directly discretizes the raw pore space into computing elements by preserving pore geometry (voxels can be used as the computing elements), and then transport equations (such as Navier–Stokes or Laplace equations) are numerically solved or approximated on the computational meshes [3, 17, 37, 107]. The lattice Boltzmann method (LBM), finite-element method (FEM), and finite volume method (FVM) are commonly used to approximate or solve transport equations at the pore scale. Generally, DNS can provide direct insight into the impact of pore microstructure on transport properties, but it has severe limitations in computational intensity. The 3D digital microstructure with large representative size and high resolution usually contains hundreds of millions (or even billions) of computational elements (or voxels). As a result, massively parallel programming, long computing time, high-performance computing (HPC) platform, and large data storage are usually required to run such large-scale numerical simulations [71, 91]. The computation-intensive nature of DNS makes it difficult to accommodate all details of pore microstructures and involve all relevant transport physics.
As discussed above, both PNM and DNS approaches have their own limitations, which have been long recognized by the DRP community. More recently, many attempts have been made to develop surrogate microstructure–property models through artificial intelligence, to rapidly and accurately predict macroscopic properties from the measurable microstructural informatics. Due to the powerful capacities in massive data analysis and hidden rule exploration, machine/deep learning algorithms are becoming increasingly popular in this field, especially the convolutional neural network (CNN). CNN [2] is capable of automatically extracting task-related features from spatial data such as images through its convolution layers, avoiding the manual feature selection procedure, and it has achieved tremendous success in the computer vision field. Therefore, many similar studies have been conducted to construct CNN-based surrogate models for permeability prediction, where the 2D or 3D digital images of porous microstructures are directly used as the input data [48, 55, 96, 97, 99, 100, 110]. Besides, CNN has also been applied to establish the linkages between microstructures and other macroscopic properties/behaviors for various heterogeneous materials, including effective thermal conductivity [108], effective elastic moduli [22, 68], effective diffusivity [109], P/S-wave velocity [56], formation factor [85], and fluid velocity filed [90].
However, despite the rapid growth in publications, the CNN-based surrogate modeling strategy is not without limitations, at least in its current forms. (1) Shortage of training images: A reliable predictive CNN model usually requires a large number of training images to feed it, but acquirement of high-quality 3D digital rocks is quite expensive, which can explain why most of the previous studies only use 2D images or sphere packing to investigate the microstructure–property relations of porous media. (2) Heavy computational burden: High computational intensity and excessive memory requirement are the inherent challenges of the 3D CNN algorithm, which strictly limit both the quality and quantity of 3D training images. Using 3D representative elementary volumes (REVs) of digital rocks to fit a CNN model usually demands an HPC platform. (3) Feature extraction problem: Kernels (convolving windows) are applied across the input image to extract local features, but the internal connection of components, as well as the relative spatial relationships, are not captured by the convolution layers of CNN. It means that the global features of porous media (such as long-distance connectivity and topological information) that are crucial to transport properties are rarely considered. (4) Rotation dependence: The internal representation of a pore microstructure in CNN is not independent of the view angle, which means that rotation of the input image can potentially affect the prediction result. This issue can be solved through data augmentation, but the computational cost of CNN model training will be dramatically increased. (5) Over-fitting problem: The CNN model is prone to over-fitting due to a large database for training. (6) Low-level interpretability: The complicated CNN architecture, formed by a deep stack of distinct layers, is often referred to as a “black box”, because it is difficult to understand the underlying mechanics and no inherent way exists to interpret how features influence a particular prediction. (7) Inflexibility: Once a CNN model is fitted for physical property prediction, both the size and resolution of input images are fixed, which is very inflexible for the common cases where adjustments of image size or resolution have to be made without losing information.
As discussed above, the poor explainability of CNN does not contribute to a good understanding of the microstructure–property linkages. In contrast, results from simple regression algorithms, such as linear regression, decision tree, random forest, support vector machine, and shallow neural network, are much easier to be interpreted, which helps to reveal the underlying mechanisms of the microstructure–property relationships. In addition, CNN is free from manual feature-extraction, but in porous media research, this feature does not constitute a comparative advantage over other regression algorithms that require predefined feature variables. This is because various morphological descriptors that quantitatively characterize pore microstructures have already been properly designed (as listed in Table 1). These descriptors can provide measurable microstructural informatics from multiple perspectives to predict macroscopic physical properties. Compared with the unreadable features extracted by CNN, morphological descriptors characterize porous microstructures from multiple perspectives with clear physical indications, and they form the feature pool that can be readily used to develop the surrogate microstructure–property relationships through simpler regression algorithms. Specifically, feature selection can be conducted to identify the morphological descriptors that are significant to permeability and remove the abundant and irrelevant ones, through which the microstructural complexity is reduced to a limited number of descriptive parameters related to permeability, and then, a high-fidelity data-driven prediction model can be achieved (as illustrated in Fig. 1). More importantly, the dependence of permeability on microstructural characteristics can also be well interpreted through the feature selection process, providing a deep insight into the microstructure–permeability relation.
Table 1.
The collected morphological descriptors for quantitative characterization of porous microstructures
| Index | Morphological descriptor (Evaluation method) |
Denotation | Data dimension | Representative references |
|---|---|---|---|---|
| D1 | Absolute porosity | 1 | Carman [21], Adler [1] | |
| D2 | Effective porosity | 1 | Géraud [39], Fu et al. [34] | |
| D3 | Specific surface area | S | 1 | Liang et al. [69], Cui et al. [29] |
| D4 | Integral of mean curvature | 1 | Lehmann et al. [66] | |
| D5 | Integral of total curvature | 1 | Vogel et al. [105] | |
| D6 |
Geometrical tortuosity (Direct shortest path search method) |
1 | Koponen et al. [61], Cecen et al. [23] | |
| D7 |
Geometrical tortuosity (Skeleton shortest path search method) |
1 | Sevostianova et al. [95], Fu et al. [35] | |
| D8 |
Constriction factor (Mercury intrusion porosimetry simulation) |
1 | Holzer et al. [47], Berg [10] | |
| D9 |
Constriction factor (Morphological opening method) |
1 | Berg [10], Dong et al. [31] | |
| D10 | Mean chord length | 1 | Coker et al. [26], Bertei et al. [12] | |
| D11 | Average pore coordination number | 1 | Hormann et al. [49] | |
| D12 | Average pore size (Continuous method) | d | 1 | Münch and Holzer [78] |
| D13 | Average pore size (Discrete method) | d | 1 | Holzer et al. [46] |
| D14 |
Average pore size (Morphological opening method) |
d | 1 | Paterson [82], Dong et al. [31] |
| D15 | Average pore size (Random point method) | d | 1 | Torquato [102] |
| D16 | Average pore size (Skeleton method) | d | 1 | Delerue et al. [30] |
| D17 | Average pore throat size | 1 | Liang et al. [69] | |
| D18 | Effective pore size | 1 | Blair et al. [15] | |
| D19 | Hydraulic pore diameter | 1 | Bear [8] | |
| D20 | Characteristic length I | 1 | Coker et al. [26] | |
| D21 | Characteristic length II | 1 | Coker et al. [26] | |
| D22 | Characteristic length III | 1 | Coker et al. [26] | |
| D23 | Characteristic length IV | 1 | Ioannidis et al. [52] | |
| D24 | Average connectivity distance | 1 | Knudby and Carrera [58] | |
| D25 | Characteristic length V | 1 | Hilfer [45] | |
| D26 | Fractal dimension | 1 | Yu and Cheng [118] | |
| D27 | Succolarity | 1 | Xia et al. [112] | |
| D28 | Lacunarity | 10 | N’Diaye et al. [79] | |
| D29 | Chord length distribution | 70 | Muche and Stoyan [77], Cui et al. [29] | |
| D30 | Lineal path function | L(z) | 50 | Hilfer [45], Cui et al. [29] |
| D31 | Spherical contact distribution function | 16 | Lehmann et al. [66] | |
| D32 | 1st Minkowski function | 16 | Vogel et al. [105] | |
| D33 | 2nd Minkowski function | 16 | Armstrong et al. [5] | |
| D34 | 3rd Minkowski function | 16 | Vogel et al. [105] | |
| D35 | 4th Minkowski function | 16 | Armstrong et al. [5] | |
| D36 | Two-point correlation function | 50 | Blair et al. [15], Fu et al. [33] | |
| D37 | Two-point cluster correlation function | 50 | Jiao et al. [53], Cui et al. [29] | |
| D38 | Normalized auto-covariance function | R(r) | 50 | Bentz and Martys [9] |
| D39 | Pair connectivity function | H(r) | 50 | Knudby and Carrera [58], Fu et al. [38] |
| D40 | Surface-surface correlation function | 20 | Rubinstein and Torquato [88] | |
| D41 | Surface-void correlation function | 50 | Rubinstein and Torquato [89] | |
| D42 | Local porosity distribution | 100 | Biswal et al. [14], Fu et al. [38] | |
| D43 | Local porosity distribution | 100 | Hilfer [45], Fu et al. [38] | |
| D44 | Local percolation probabilities | 50 | Cosenza et al. [27], Fu et al. [38] | |
| D45 | Local percolation probabilities | 50 | Cosenza et al. [27] | |
| D46 | Total fraction of percolating cells | 75 | Latief et al. [65], Fu et al. [38] | |
| D47 | Total fraction of percolating cells | 75 | Hilfer [45], Fu et al. [33] | |
| D48 | Pore coordination number distribution | 20 | Hormann et al. [49] | |
| D49 |
Pore size distribution (Continuous method) |
p(d) | 50 | Münch and Holzer [78] |
| D50 |
Pore size distribution (Discrete method) |
p(d) | 30 | Holzer et al. [46] |
| D51 |
Pore size distribution (Morphological opening method) |
p(d) | 20 | Dong et al. [31] |
| D52 |
Pore size distribution (Random point method) |
p(d) | 30 | Torquato [102] |
| D53 | Pore size distribution (Skeleton method) | p(d) | 30 | Delerue et al. [30] |
| D54 | Pore throat size distribution | 15 | Lindquist et al. [70] | |
| D55 | Coarseness | C(L) | 100 | Quintanilla and Torquato [84] |
Fig. 1.
The data-driven framework to investigate the microstructure–permeability relation for natural porous rocks, and it contains six functional modules: (1) Digital rock acquirement, (2) Stochastic microstructure reconstruction, (3) Quantitative microstructure characterization, (4) Pore-scale flow simulation, (5) Feature selection, and (6) Data-driven modeling (the paired observations are comprised of the morphological descriptors selected from the 5th module and the permeability values evaluated from the 4th module, based on which data-driven machine learning models can be trained to construct the nonlinear microstructure–permeability mapping)
Although simple regression algorithms have been adopted to model physical properties of porous media in the previous studies [32, 87, 101, 104, 106], they provide little insightful understanding of the linkages between macroscopic physical properties and microstructural characteristics, and the corresponding pore-scale behaviors are still poorly understood. This study distinguishes itself from previous studies in the following five aspects (as graphically illustrated in Fig. 1). (1) Plenty of 3D digital rocks with diverse morphologies are acquired from micro-CT scanners at high resolution, which are used to construct the predictive model with strong generalization capacity. (2) A large number of 3D microstructure samples are stochastically reconstructed by preserving statistical equivalence, morphological similarity, and transport properties, which are used as the raw data to capture the stochasticity in permeability modeling. (3) A wide variety of morphological descriptors are collected through an extensive literature study, aiming to provide comprehensive characterization of porous microstructures. (4) High-fidelity simulations of pore-scale fluid flow passing through the REVs of digital microstructures are performed to obtain reliable permeability values. (5) Feature selection is conducted to identify the optimal feature set that best represents permeability, based on which a data-driven model with excellent prediction performance can be constructed, and the model interpretability can be also enhanced.
In summary, this work proposes a data-driven framework to investigate the dependence of macroscopic physical properties on microstructural characteristics for porous media. Here, we focus on the intrinsic permeability of porous rocks, but this framework is generally applicable to study other physical properties, such as hydraulic, thermal, electrical, diffusional, and mechanical behaviors. The remainder of this paper is organized as follows. The methodology of the data-driven framework is explained in detail in Sect. 2, and the raw datasets are also prepared and organized, including microstructure sample generation, morphological descriptor extraction, and permeability evaluation. In Sect. 3, different types of feature selection are tested to identify the morphological descriptors that are significant to permeability. In Sect. 4, the optimal feature sets identified by the wrapper methods used to construct data-driven prediction models, and regression performances are also deeply analyzed. The data-driven models are compared with two popular empirical/analytical formulas in terms of prediction accuracy and generalization performance in Sect. 5. Finally, the key findings and relevant thinking are discussed, and the main contributions of this work are also summarized in Sect. 6.
Methodology and data preparation
As illustrated in Fig. 1, the proposed data-driven framework comprises six functional modules: digital rock acquirement, stochastic microstructure reconstruction, quantitative microstructure characterization, pore-scale flow simulation, feature selection, and data-driven modeling. These modules are integrated together to form a data-driven approach to investigating the microstructure–permeability relation of natural porous rocks. The first four modules are briefly explained in this section, while the last two modules will be detailed in Sect. 3 and Sect. 4, respectively.
Digital rock acquirement
To investigate how microstructural characteristics determine intrinsic permeability, a variety of porous media (mainly porous rocks) are used in this study, including sandstones, carbonate rocks, sand packs, and synthetic silicas among others. The pore network systems inside these porous media are rather different in terms of their geometry, topology, fractal property, and statistical attribute, which assures that the resulting prediction model holds for a diverse range of porous media. For sedimentary rocks in hydrocarbon reservoirs, the porosity values generally vary from to in sandstones and from to in carbonates. As illustrated in Fig. 2, the porous media samples used in this study have a wide porosity range varying from to . Besides, the permeability values also broadly vary across 4 orders of magnitude, as shown in Fig. 8.
Fig. 2.

The porosity distribution of the porous media samples used in this study
Fig. 8.
The permeability results evaluated from lattice Boltzmann simulations for the digital microstructure dataset with 1455 samples
Modern microscopy imaging techniques can be used to characterize the internal geometries of opaque porous rocks at the micro-scale. Here, 3D digital rock samples are acquired from micro-CT scanning, and they can be used for subsequent studies including microstructural analyses and pore-scale numerical simulations. The raw micro-CT image is usually in grayscale, as shown in Fig. 3a. It is necessary to convert the raw grayscale image to a segmented format that permits quantitative characterization of the porous microstructure and pore-scale simulation of fluid flow. As illustrated in Fig. 3, the raw micro-CT image of a Mt. Simon sandstone sample [59] is denoised and segmented. The binary segmentation is often referred to as the digital microstructure, where the pore space is separated from the solid matrix. The digital microstructure provides a computational mesh for quantitative characterization of the pore network system and numerical simulation of pore-scale flow. More details on image processing and segmentation can be found in relevant references [50, 93].
Fig. 3.
Illustration of image processing and segmentation: (a) The raw micro-CT image of a Mt. Simon sandstone sample (resolution is 2.80 m, and image size is voxels); (b) the grayscale image after denoising and enhancement; (c) the histogram of voxel grayscale value; and (d) the binary image segmented by a global thresholding method (pore space is shown in white, and solid matrix is shown in black)
It is noted that the micro-CT images used here are collected from several open-access databases, and these images are processed and segmented using ImageJ [51], a popular image processing tool in the DRP community. Due to the high costs of rock core drilling and microscopy imaging, there are only a limited number of 3D digital rock samples available. A total of 185 micro-CT images are used in this study, covering 37 types of porous media with distinct morphological features (the representatives of them are shown in Figs. 7 and 9).
Fig. 7.
Evaluation of intrinsic permeability through lattice Boltzmann simulation: (a) the 3D digital microstructure of a Mt. Simon sandstone sample; (b) the boundary conditions; and (c) the steady-state fluid velocity field inside the porous medium
Fig. 9.
Digital rock samples, image segmentation, and lattice Boltzmann simulations: The micro-CT scanning images of (a) Ketton carbonate [92], (d) Fontainebleau sandstone [65], (g) Savonnières carbonate [20], and (j) Leopard sandstone [44]; (b), (e), (h), and (k) are the segmented images; (c), (f), (i), and (l) are the steady-state flow velocity fields inside porous microstructures
Stochastic microstructure reconstruction
The transport properties of porous media usually exhibit strong uncertainty, due to the random distribution of pore bodies. As a result, the limited number of digital rock samples obtained from micro-CT scans is far from sufficient to cover all possible morphology configurations of pore microstructures. In general, the complete computational dataset [33, 38] is an ensemble of representative/statistical volume elements that cover all morphological possibilities and share the same averaged characteristics, based on which a generalized prediction model can be achieved with high reliability.
Stochastic microstructure reconstruction [18, 33, 81] is an effective and economical approach to generating statistically equivalent samples of porous media, and the numerous reconstructed samples can be used to investigate the microstructure–property correlations when the availability of real porous media samples is limited. In this work, a high-fidelity reconstruction method developed in our previous study [33, 36, 38] is adopted to generate 3D pore microstructure samples. This method first characterizes the morphology patterns of the real 3D microstructures by fitting statistics-informed neural networks, based on which virtual 3D microstructure samples can then be generated via probability sampling. These virtual samples have been proven to preserve statistical equivalence, morphological similarity, long-distance connectivity, and transport properties of the real ones, and more details can be found in relevant references [33, 36, 38]. As shown in Fig. 4, Fontainebleau sandstone samples with different porosities are taken as examples to illustrate stochastic microstructure reconstruction using this new method. Guided by the morphological information extracted from the real digital microstructures, a total of 1270 virtual microstructure samples are reconstructed, with the image size varying from to voxels. The scanned digital rocks together with the reconstructed microstructure samples compose the raw dataset of 1455 samples in total for subsequent analyses.
Fig. 4.
Stochastic microstructure reconstruction (image size: voxels): (a–c) are the scanned microstructures of Fontainebleau sandstones with different porosities ; (d–f) are the representative reconstructed microstructures
Quantitative microstructure characterization
Quantitative characterization of porous microstructures in an explicit expression is the essential prerequisite to exploring the microstructure–property linkage of porous media. The pore space inside natural porous rocks usually exhibits great disorder and strong randomness, which needs to be quantified in statistical terms. Through quantitative characterization [103, 114], the microstructural complexity of a porous medium can be reduced to a small set of morphological descriptors related to the macroscopic physical property of interest. A broad range of microstructure characterization approaches have been developed for porous media, such as statistical characterization, geometrical measurement, topological representation, and fractal analysis. As listed in Table 1, the commonly used morphological descriptors are collected through an extensive literature study, and they will be used as the microstructural features for permeability prediction. Many of these descriptors have been used to investigate the microstructure–property relations of porous media in the previous studies.
These morphological descriptors characterize porous microstructures from different perspectives, and they can be roughly grouped into four levels. Porosity and specific surface area are the typical descriptors at the first level to simply represent the global/mean properties of porous microstructures via single numbers, but they ignore the detailed morphological features of pore networks that may have significant effects on transport processes. As to the second level, local or size-dependent features are measured by such morphological descriptors as local porosity distribution, coarseness, local percolation probabilities, and lacunarity. When it comes to the third level, geometric attributes of porous media are quantified from various aspects such as pore shape, pore size, and surface roughness. The frequently used descriptors includes pore/throat size distribution, mean curvature, chord length distribution, lineal path function, and spatial correlations functions. The fourth level focuses on the topological characteristics of pore microstructures, which is related to long-distance connectivity and percolation of pore networks. Total curvature (Euler characteristic), two-point cluster correlation function, pair connectivity function, total fraction of percolating cells, and succolarity are commonly used indicators of connectivity. Pore coordination number represents the number of adjacent pore bodies connected to a specific pore. Besides, geometrical tortuosity characterizes the sinuosity and complexity of percolation paths inside porous media, while constriction factor quantitatively represents cross-sectional variation along pore channels.
All morphological descriptors in Table 1 are extracted from the microstructure dataset with 1455 samples, and they serve as the possible predictors to construct data-driven models for permeability prediction, as illustrated in Fig. 1. It is noted that some descriptors have multiple definitions and therefore different evaluation methods, and they are all used in this study to achieve a microstructure characterization as comprehensive as possible. As shown in Fig. 5, the results of average pore size, geometrical tortuosity, and construction factor are computed using different methods. Generally, different evaluation methods yield inconsistent values of morphological descriptors, but the results show similar changing trends and are highly correlated as well. There is no general standard to judge the rationality of the descriptor result calculated from a specific evaluation method, so feature selection could be an effective mean to choose an appropriate evaluation of a morphological descriptor. Besides, the evaluation results of another 12 representative descriptors are provided in Fig. 6.
Fig. 5.
The morphological descriptors with multiple definitions extracted from the digital microstructure dataset with 1455 samples
Fig. 6.
The results of representative morphological descriptors extracted from the digital microstructure dataset with 1455 samples
As mentioned above, the digital rocks are collected from several open-access databases, so the image resolutions (voxel sizes) of them are slightly different, which are all around 5 m. It means that microstructural analyses are conducted in voxel domains with different length scales to compute morphological descriptors. For the dimensionless descriptors, such as porosity, geometrical tortuosity, constriction factor, and poor coordination number, no additional data processing is required. As to the descriptors with length dimension, such as specific surface area, mean curvature, average pore size, and characteristic length, they are all quantified using voxel as the basic length unit, instead of converting them into the physical length scale. This treatment enables the seamless combination between the morphological descriptors in the voxel length unit and the LBM permeability in the lattice length unit, just by setting the lattice length equal to the voxel size for each porous media sample.
Pore-scale flow simulation
For pore-scale simulation of fluid flow, lattice Boltzmann method (LBM) [62] is more mathematically rigorous than pore network modeling (PNM), and the former can also provide more reliable permeability evaluations for porous media with complicated geometries [113]. Besides, the LBM simulation is directly performed on voxel domain of digital microstructures without any simplification, and the computed permeability values in the lattice unit can be directly linked to the morphological descriptors in the voxel unit, avoiding additional data conversion/processing. Therefore, LBM is adopted in this work to evaluate the intrinsic permeability values of these digital microstructure samples.
Basic theory of LBM
LBM [62, 111, 119] models the fluid flow through a time-dependent distribution of fluid particles propagating on a regular lattice. In DRP research, pore voxels in digital rock images serve as the regular lattice for LBM to simulate pore-scale fluid flow, and each lattice node is located in the center of corresponding pore voxel. The numerical grid of lattice Boltzmann simulation completely coincides with the image voxel grid in this study. The particle distribution function represents the probability of finding a fluid particle with the lattice velocity in the location and at the time t. Beginning with an initial state, moves from one lattice node to its neighboring nodes at each time step, and evolves itself locally subject to both mass and momentum conservation.
The conventional LBM scheme with the D3Q19 lattice arrangement and the Bhatnagar–Gross–Krook (BGK) collision operator [24] are adopted in this study. The evolution of along the direction of from the time t to can be expressed as
| 1 |
where is the single-relaxation time, is the equilibrium distribution function, and the subscript i indicates the direction of lattice velocity around the lattice node. The relaxation time is a function of kinematic lattice viscosity of simulated fluid, i.e., , where is the lattice speed of sound and it is assigned with the dimensionless value of .
The equilibrium distribution function corresponds to an ideal state where the particle distributions tend to a specific macroscopic state, to recover the macroscopic Navier–stokes equation. For the D3Q19 lattice arrangement with BGK collision operator, is expressed as [24]
| 2 |
where is the weight factor of D3Q19 lattice structure, is the fluid density, and is the macroscopic fluid velocity. For the D3Q19 lattice model, the weight factors are equal to , , and for the velocity directions of the central lattice node, face-connected neighbors, and edge-connected neighbors, respectively.
At the end of each time step, the macroscopic properties of fluid flow, including density and velocity , can be approximated from through the following equations, and these macroscopic properties will be used for the LBM computation at the next time step
| 3 |
| 4 |
where n is the number of lattice directions ( in D3Q19 lattice structure used in this study).
Permeability evaluation
Driven by a constant pressure difference between the inlet and outlet faces, LBM is performed on the cubic digital rock sample to simulate a single-phase fluid flow with low Reynolds number () passing through it. In this study, small pressure gradients are applied to 3D porous media samples, to ensure permeability results are evaluated from laminar fluid flows. As shown in Fig. 7, the Mt. Simon sandstone sample in Fig. 3 is taken as the example to illustrate lattice Boltzmann simulation of pore-scale fluid flow. When the fluid flow reaches a steady state, it can be described by Darcy’s law, and the intrinsic permeability of this sample is quantified by the following equation:
| 5 |
where denotes the pressure gradient along the direction of macroscopic fluid flow, is the dynamic viscosity, and denotes the volume averaged fluid velocity across the entire simulation domain.
As the initial condition of lattice Boltzmann simulation is less important for steady-state flows and corresponding long-term behaviors, we simply assign the initial velocity and initial density to all lattice nodes in the simulation domain [64]. Three types of boundary conditions are adopted for the pore-scale simulation: the no-slip boundary condition on the pore-solid surface, fixed pressure boundary condition at the inlet and outlet faces, and periodic boundary condition applied to the surfaces that are parallel to the main flow direction. The complete bounce-back scheme is implemented to simulate the no-slip boundary condition, where a fluid particle bounces back to the node it comes from with no relaxation when it meets a solid node. To apply the constant pressure difference, two void layers are added to both inlet and outlet faces [34, 54], and the pressure difference between the inlet and outlet is expressed by fluid density difference. The lattice Boltzmann simulation runs continuously until reaching the user-prescribed convergence criterion. In this case, the fluid flow is assumed to be stable when the standard deviation of average kinetic energy falls below (the maximum number of iterations is 60,000). As shown in Figs. 7 and 9, lattice Boltzmann simulations are performed on eight representative digital rock samples to achieve the steady-state flow velocity fields for permeability evaluation.
The permeability value computed from lattice Boltzmann simulation is in dimensionless lattice unit, and it can be converted to the physical unit via the following equation [98]:
| 6 |
where and are the permeability values in the physical and lattice unit, respectively; and and are the lengths of any identical feature in the physical sample and the LBM domain, respectively. As the numerical grid of LBM coincides with the voxel grid of the digital microstructure in this study, the value of is equal to the image resolution (voxel size).
For each porous media sample, lattice Boltzmann simulations of fluid flow are conducted along three-axial directions, and the average value of three directional permeabilities is used to investigate the microstructure–permeability relationship. As recorded in Fig. 8, the permeability values of the digital microstructure dataset with 1455 samples are plotted both in lattice unit and physical unit. From the above figures, one can see that the permeability values span in a broad range over 4 orders of magnitude. To avoid extra data processing, the permeability values in lattice unit will be used to explore the microstructure–permeability relation via feature selection and data-driven modeling, as illustrated in Fig. 1.
Feature selection
As listed in Table 1, a variety of morphological descriptors that quantitatively characterize the internal microstructures of porous media are collected. However, these descriptors are not equally important for permeability, and some of them represent overlapping features. In addition, the inconsistent results of morphological descriptors obtained from different evaluation methods can also negatively affect investigation of the microstructure–permeability linkage. Hence, a brute-force regression model based on all available morphological descriptors cannot provide accurate and reliable permeability prediction, due to the noise from irrelevant (or less important) descriptors and the conflicts between overlapping descriptors. Besides, the unnecessary involvement of irrelevant and abundant features can increase the model complexity and make it harder to interpret. Therefore, feature selection is an indispensable step for predictive model construction, where the most relevant and significant features are to be identified from a large set of morphological descriptors in Table 1.
The objectives of feature selection in this work include: (1) enhancing interpretability of the implicit regression model to obtain deep insights into the underlying dependence of permeability on microstructural characteristics; (2) reducing the computational complexity and avoiding over-fitting to built a cost-effective predictor using the selected features; (3) achieving a generalized and rational model with the optimal performance in permeability prediction.
For a dataset of m observations consisting of n input feature variables and an output permeability value , various methods can be applied to select the feature variables that are important to the response . Generally, feature selection techniques [42, 67] can be divided into three categories: filter, embedded, and wrapper methods. Among them, filter type feature selection is independent of learning algorithms, while wrapper and embedded methods interact with a particular learning process. All these methods are tried and tested in this study to identify the most suitable feature selection for the morphological descriptors that best represent permeability of porous media.
Filter type feature selection
Filter type feature selection [42, 67] assesses feature importance according to certain data characteristics, so it is unrelated to any learning algorithms. Typically, a filter method consists of two steps: feature importance ranking and feature filtering. Different feature evaluation criteria have been proposed to rank feature importance, such as feature correlation, mutual information, the feature discriminative ability, the feature ability to maintain the data manifold, and the feature capacity to reconstruct the raw data. Four representative criteria of feature importance evaluation are covered in this study, which are Pearson’s correlation coefficient , RReliefF importance weight [86], F-test importance score F [6], and nearest-neighbor-based feature weight [116]. More algorithm details can be found in relevant references.
Embedded type feature selection
Embedded methods [42, 67] conduct feature selection during the learning processes, which are deeply embedded in specific learning algorithms. For example, during the training process of a decision tree [72], feature importance is evaluated from the sum of changes in the mean squared error due to splits on each feature and the number of branch nodes. For a random forest [19], feature importance can be evaluated by permutation to measure the influence degree of a feature variable in predicting the response. As to Gaussian process regression [94], feature importance can be evaluated from corresponding separate length scales of the kernel function. In this study, the above three learning algorithms are tested to assess the importance of morphological descriptors to permeability modeling, and more algorithm details can be found in relevant references.
Wrapper type feature selection
Wrapper type feature selection is more applicable to heterogeneous features, compared to the filter and embedded methods. Considering the dimension differences between morphological descriptors listed in Table 1, wrapper type feature selection is attractive. Wrapper methods [43, 60] assess the quality of feature selection according to the prediction performance of the predefined learning algorithms. It searches the optimal feature subset through greedily evaluating the possible combinations of features based on a certain evaluation criterion. For regression problems, the coefficient of determination , the mean-squared error (MSE), and the mean-absolute error (MAE) can be used as the metrics to evaluate the model performance, which are mathematically expressed by the following equations, respectively:
| 7 |
| 8 |
| 9 |
where and are the target and predicted permeability respectively corresponding to the i-th porous media sample, and is the average value of the target permeability. Among them, quantifies the degree to which the feature variables explain the variation of the response, and its value ranges from 0 to 1, where a larger value indicates a better model performance.
Exhaustive search is a “brute-force” strategy in wrapper type feature selection, which usually requires enormous amounts of computation, especially when the number of feature variables is large. By contrast, greedy search strategies are of lower computation cost, which can be further divided into two categories: forward selection and backward elimination. Here, the wrapper method with sequential forward adding strategy [40] is adopted, as explained in Fig. 10.
Fig. 10.

The flowchart of wrapper type feature selection through a sequential forward adding strategy (it should be noted that this flowchart is only for one round of feature selection, and the remaining rounds just repeat this procedure)
Starting with a null model, each morphological descriptor in Table 1 is used individually to construct a predictive surrogate model, and the descriptor that achieves the best predictive performance (the maximum or the minimum RMSE value) is picked out as the first selected feature. A new predictive model with two features is then constructed by sequentially combining the previously selected feature with one of the remaining descriptors, and the descriptor resulting in the largest or the smallest RMSE is selected as the second feature. The above procedure is repeated iteratively until no improvement of prediction performance or reaching the desired number of included features, and a subset of features are consequently selected through this performance-orientated process.
Feature selection results
As listed in Table 1, the first 27 morphological descriptors are in the format of a single number, while the remaining 28 descriptors are in the form of distributions with different data dimensions. Generally, the filter and embedded methods are not applicable to feature selection with heterogeneous data, while the wrapper type feature selection possesses good versatility. Considering the above situation, feature selection is first performed on the first 27 morphological descriptors using different methods, and later the wrapper method is applied to select features from the entire feature pool with all 55 descriptors.
Data normalization
After microstructural characterization and pore-scale simulation, the paired data with m observations can be obtained to study the microstructure–permeability relation. For both feature selection and data-driven modeling, the scale of labeled data can greatly affect the results, and thus, data normalization is required to deal with this issue. As the jth feature variable, within a range of interest can be scaled using the minimum and maximum values, given by
| 10 |
where and are the maximum and minimum values, respectively, of the jth feature variable. As to the output variable, the range of permeability value is not known for unseen data, so it is statistically normalized as follows:
| 11 |
where and are the mean value and standard deviation, respectively, of permeability data.
Filter type feature selection results
As plotted in Fig. 11, the results of feature importance ranking are estimated for the first 27 descriptors using four different filter methods. Due to the evaluation criterion difference of feature selection, the importance ranking results estimated from these methods are not completely consistent with each other, but the overall assessment results are similar. In all four filter methods, effective porosity (D2) and average connectivity distance (D24) are identified as the influential microstructure characteristics to permeability (), which agrees with the common knowledge of porous media [1, 21, 34].
Fig. 11.
The feature importance ranking results estimated from four filter methods
However, specific surface area (D3) is evaluated to be an insignificant/irrelevant feature, which is contrary to the general consensus that the specific surface area is critical to permeability of porous media [1, 21, 34]. Filter methods evaluate the importance of feature variables individually, but a feature variable that is recognized to be unimportant by itself can be significant to the response when used with the other features [42]. Basically, filter methods are unable to detect the joint importance of multi-variable features, which is one of their main drawbacks.
Embedded type feature selection results
Three embedded methods are also used to assess the importances of the first 27 morphological descriptors to permeability, and corresponding results are plotted in Fig. 12. Because different learning algorithms are embedded in these three feature selection processes, the importance rankings of morphological descriptors are not completely consistent. Similar to the results of filter methods, effective porosity (D2) and average connectivity distance (D24) are also selected as important microstructure features by embedded methods, but specific surface area (D3) is assigned with low importance scores, especially in the regression tree and the GPR model. Besides, embedded type feature selection is associated with specific learning algorithms, which is inflexible for prediction model construction.
Fig. 12.
The results of feature importance ranking evaluated from three embedded methods
In summary, the intended purpose of feature selection has not been achieved using the filter or embedded method. Feature importance has been missed for some specific descriptors that are known to be critical to permeability of porous rocks, while only the scaler-valued morphological descriptors in Table 1 are covered by these two methods. This task will be continued with the wrapper type feature selection in the following part.
Wrapper type feature selection results
As explained in Sect. 2.5.3, wrapper type feature selection is highly interrelated to the learning algorithm. Therefore, it is crucial to choose an appropriate learning algorithm for both wrapper-based feature selection and data-driven modeling. On the one hand, the chosen learning algorithm should possess strong learning capacity to deal with high-dimensional data; on the other hand, the model response should also be sensitive to influential feature variables to capture the feature significance. After conducting the comparison between linear regression, decision tree, random forest, support vector machine, and feed-forward neural network (FNN), the FNN with a shallow architecture is found to be the most appropriate learning algorithm for this study.
Feed-forward neural network
Artificial neural networks [117] are function approximators to map the inputs to the output through many interconnected computation elements called neurons. Each elementary neuron possesses a certain degree of approximation capacity, and a powerful learning performance can be achieved by cohesively combining many neurons. It has been proved that a fairly simple neural network is capable of fitting many practical functions [63]. The feed-forward neural network with a shallow architecture is adopted to construct the implicit microstructure–permeability model in this study.
As illustrated in Fig. 13, morphological descriptors are used as the feature variables to feed an FNN model with one or two hidden layer(s), and the final output is a permeability prediction. The predicted permeability is computed through a series of forward-propagation equations that occur at particular layers, given by
| 12 |
where denotes the input features; denotes the output of the kth layer; and are the weight matrix and bias of the ith layer, respectively; denotes the activation function (hyperbolic tangent function is adopted here).
Fig. 13.
The graphic illustration of an FNN model with two hidden layers for permeability prediction
Essentially, the above FNN model is a vector-valued network surrogate to approximate the input–output relation of the - mapping, and the approximation function can be mathematically expressed as follows:
| 13 |
where denotes the approximation function of the FNN model; and are the data dimensions of the input and the output, respectively.
The next key issue here is to optimally adjust the weight matrices and bias vector of the neural network by making full use of the available labeled data. Basically, data-driven training is to optimize and of the neural network by minimizing of the discrepancy between the targets and the outputs for the observational data. This optimization problem can be mathematically expressed as follows:
| 14 |
where denotes the loss function, and is the weight regulation constant. The first term of the loss function is the mean-squared error (MSE) to represent the discrepancy between the targets and the predictions . The second term is the weight regulation term, also called weight decay, which can force the network response to be smoother and thus to reduce over-fitting.
The above minimization problem can be solved by many optimization algorithms, such as stochastic gradient descent methods. Generally, Levenberg–Marquardt algorithm (LMA) [76] is considered to be the most cost-effective method to train moderate-sized feed-forward neural networks (up to several hundred weights) with high accuracy, especially for regression problems. Therefore, LMA is adopted here to obtain the optimal weights and biases of feed-forward neural networks for the purpose of function approximation. Besides, cross-validation is usually performed to avoid over-fitting, thereby improving the generalized predictive ability for new observations. More details about parameter optimization and cross-validation of artificial neural networks can be found in relevant references [63, 117].
Feature importance indicator
After data normalization (as explained in 3.1), the entire dataset is randomly split into three subsets: training (50%), validation (25%), and test (25%). The training dataset is used to fit the neuron network, where network parameters are optimally adjusted to minimize regression error. The validation dataset is used to measure network generalization, and the training process is halted when generalization stops improving, so as to avoid over-fitting. The testing dataset is used to provide an independent measure of the model performance on unseen data.
Due to the randomness of initial configurations (such as random sampling of initial network parameters, and random division of the entire dataset for training, validation, and testing), training the neural network multiple times usually generates different results (as illustrated in Fig. 14a). Here, the feed-forward neural network is trained for 100 times for each case of input features, and the average value of over these 100 trials is used as the indicator to represent feature importance for feature selection using the wrapper method. The hyper-parameters of feed-forward neuron networks are summarized in Table 2.
Fig. 14.
Regression performances of the FNN models trained by individual features: (a) The varying values of for different trails; (b) The average regression performance is used as the indicator to represent feature importance for all 55 morphological descriptors in Table 1
Table 2.
The hyper-parameters of neural networks trained by using Levenberg–Marquardt algorithm
| Data-driven modeling |
Model hyper-parameters | Algorithm hyper-parameters | |||||||
|---|---|---|---|---|---|---|---|---|---|
| Number of hidden layers |
Neurons in per hidden layer |
Initial damping factor |
Decrease factor for |
Increase factor for |
Minimum value for |
Maximum value for |
Regulation constant |
Maximum validation failures |
|
|
Predictive model I (Selection strategy I) |
2 | 15 | 0.1 | 10 | 20 | ||||
|
Predictive model II (Selection strategy II) |
2 | 30 | 0.1 | 10 | 50 | ||||
Note: is the combination coefficient in Levenberg–Marquardt algorithm, and its meaning can be found in the appendix section of this paper
In Fig. 14b, the average values of over 100 trials are used as the indicators to represent feature importance for all 55 morphological descriptors in Table 1, where each descriptor is used individually to train the neural network. The first 27 morphological descriptors in Table 1 are in the format of a single number, which are represented by the blue bars, and we call them for convenience. As to the remaining 28 morphological descriptors, they are in the format of a distribution, which are called here.
Feature selection strategy I
In this part, feature selection is restricted to the (the first 27 descriptors in Table 1) using the wrapper method, aiming to built a cost-effective surrogate model with fewer input variables. The methodology of wrapper type feature selection is graphically illustrated in Fig. 10. In the first round of feature selection, the blue descriptors are used individually to fit the neural network one by one, and the regression performances are recorded in Fig. 14b. Effective porosity (D2) is selected as the most important feature in this round, because it yields a regression model with the best predictive performance () among all the blue descriptors. This feature selection result is consistent with that of the filter and embedded methods, as shown in Figs. 11 and 12.
In the second round, D2 is combined with the remaining descriptors one by one to jointly fit the neural network, and the regression performances are remarkably improved, as shown in Fig. 15a. The maximum value of is 0.9838, and the corresponding descriptor D3 (specific surface area) is selected as the second feature. In contrast to the filter and embedded methods, the significance of specific surface area to permeability can be well recognized by the wrapper type feature selection. Repeating the above procedures, morphological descriptors D6, D26, D1, D11, D8, D10, D17, and D18 are then successively identified, as illustrated in Fig. 15b–i.
Fig. 15.
The selected descriptor at each round via the wrapper type feature selection
As illustrated in Fig. 16a, the regression performance increases continuously when more selected morphological descriptors are included in the FNN model. However, the regression performance reaches the peak () when the first eight selected descriptors are used to train the FNN model, which is highlighted by the red star in Fig. 16a. Continuing to add more selected descriptors for model training, the regression performance starts to decline. Therefore, the optimal feature set obtained from the wrapper method contains eight morphological descriptors, which are: effective porosity (D2), specific surface area (D3), geometrical tortuosity (D6), fractal dimension (D26), absolute porosity (D1), pore coordination number (D11), constriction factor (D8), and mean chord length (D10).
Fig. 16.
The regression performance of FNN models varying with the number of selected descriptors
The above feature selection result agrees with the existing knowledge of porous media, and the selected morphological descriptors have all been directly used to build analytical/empirical formulas for permeability evaluation in the previous studies. For instance, the first three selected descriptors (effective porosity, specific surface area, and geometrical tortuosity) are used in the well-known Kozeny–Carmon relation [21, 25, 34]. Different from the filter and embedded methods that only analyze the relationship between an individual descriptor and permeability, the wrapper method selects features in a multi-variable analytical manner.
Essentially, the optimal feature set consist of eight morphological descriptors quantitatively characterize pore network systems inside porous media from seven different perspectives, based on which the dependence of permeability on microstructural characteristics can be interpreted as follows: (1) Absolute and connected porosity represent the entire pore space and the permeable portion permitting fluid to flow through, respectively; (2) Specific surface area approximately reflects the area of fluid–solid interface that provides adhesive friction to fluid flow; (3) Geometrical tortuosity measures the sinuosity of percolating pore paths that extends the average length of flow streamlines; (4) Fractal dimension is a measurement of scaling irregularity and complexity of porous microstructures in which transport phenomenon of fluid flow occurs; (5) Average pore coordination number characterizes the topology properties of porous media, which represents the number of adjacent pore bodies connected to a specific pore body; (6) Constriction factor quantifies the degrees of cross-section variation along pore channels, which can converge and diverge fluid streamlines and thus hinder transport flow; (7) Mean chord length measures the spatial distances between opposite walls of pore channels allowing fluid to pass through.
Feature selection strategy II
Despite the encouraging results, the have not been included in the above discussion. In this part, the wrapper type feature selection is performed on all morphological descriptors listed in Table 1, expecting to identify a set of descriptors that comprehensively characterizes porous microstructures, from which a deep insight into the microstructure–permeability relation can be obtained.
In the first round of feature selection, D47 (total fraction of percolating cells) is picked out, because it yields the neural network model with the best performance (), as shown in Fig. 14b. In the second round, D47 is combined with the remaining descriptors one by one to jointly fit the neural network, and the regression performances are shown in Fig. 17a. The maximum value of is 0.9921, and the corresponding descriptor D6 (geometrical tortuosity) is selected as the second feature. Repeating the above procedures, morphological descriptors D3, D2, D27, D25, D1, D11, D15, and D26 are then successively picked out, as demonstrated in Fig. 17b–i.
Fig. 17.
The selected descriptor at each round via the wrapper type feature selection
As illustrated in Fig. 16b, the regression performance of the FNN model reaches its peak (the red start) with , when the first nine selected descriptors are used as the input feature variables, Therefore, the optimal feature set contains nine morphological descriptors, which are: total fractional of percolating cells (D47), geometrical tortuosity (D6), specific surface area (D3), effective porosity (D2), succolarity (D27), characteristic length (D25), absolute porosity (D1), average coordination number (D11), and average pore size (D15). Only one green descriptor is contained in the optimal feature set, and the remainders are blue descriptors.
Comparing the feature selection results in Sect. 3.4.3 and Sect. 3.4.4, there are five morphological descriptors in common, which are D6, D3, D2, D1, and D11. Besides, the newly selected descriptor D15 quantitatively characterizes the spatial size of pore channels allowing fluid percolation, which is conceptually similar to the formerly selected descriptor D10 (mean chord length) in Sect. 3.4.3. The remaining selected descriptors still contain D47 (total fractional of percolating cells), D27 (succolarity), and D25 (characteristic length), and they provide a new perspective to understand the microstructure–permeability relation. In essence, D47, D27, and D25 are quantitative indicators to characterize the percolation degree and long-distance connectivity of the pore networks that allow fluid flow to pass through.
Data-driven modeling
In this section, the feature selection results in Sect. 3.4.3 and Sect. 3.4.4 are separately applied to construct two surrogate models to approximate the microstructure–permeability mapping through data-driven modeling. As shown in Table 3, the morphological descriptors used to construct predictive models are clearly listed. The feed-forward neural networks used for function approximation are completely same to the one used for feature selection, and the hyper-parameters are summarized in Table 2.
Table 3.
Comparisons between different predictive models in terms of permeability evaluation accuracy
| Predictive model | Involved morphological descriptors | High-permeable rocks () |
Low-permeable rocks () |
||
|---|---|---|---|---|---|
| Range of relative evaluation errors | Average error magnitude (%) | Range of relative evaluation errors | Average error magnitude (%) | ||
| Data-driven model I |
D2, D3, D6, D26, D1, D11, D8 and D10 |
||||
| Data-driven model II |
D47, D6, D3, D2, D27, D25, D1, D11 and D15 |
||||
| Kozeny–Carman relation | D2, D3 and D6 | ||||
| Berg’s relation | D2, D6, D8 and D15 | ||||
Data-driven model I
As explained in Sect. 3.4.3, the optimal feature set containing eight blue descriptors has been obtained through the wrapper type feature selection, and they are listed in Table 3. These eight morphological descriptors are used as the feature variables to fit an FNN for permeability prediction. In fact, such a predictive model has already been obtained together with the feature selection result in Sect. 3.4.3. The critical issue is how well the microstructure–permeability relation is represented by the surrogate model, which should be further analyzed.
As shown in Fig. 18a, the loss function of the neural network is quickly minimized, and the best training state is determined by the validation performance. Once the neural network is properly trained, it can be used for permeability prediction, and the results are plotted in Fig. 18b. By overall comparison, the permeability prediction results agree well with the lattice Boltzmann simulation results (the targets), especially for large permeability values. However, a clear trend can be seen from Fig. 18b is that the prediction–target discrepancy increases as the permeability value declines.
Fig. 18.
The data-driven FNN model I for permeability prediction
To make a thorough analysis of the prediction results, the entire observational data are divided into two subsets by a permeability threshold equal to 1,000 millidarcy (md). As shown in Fig. 18b, the data points in the green ellipse correspond to high-permeable rock samples, and the remaining points in the orange ellipse represent low-permeable rock samples. Obviously, the fitted FNN model is of high accuracy for permeability evaluation of high-permeable rock samples, and the evaluation errors are summarized in Table 3. As illustrated in Fig. 19a, the relative prediction error varies from 49.11% to 47.58%, with the average error magnitude as low as 7.70%. Besides, for 85.46% of high-permeable rock samples, this predictive model is able to provide permeability values with very low relative errors within ±10%.
Fig. 19.
Relative error distributions of the permeability values evaluated from the data-driven model I
However, the fitted FNN model becomes less accurate, when it comes to low-permeable porous media samples. As illustrated in Fig. 19b, the relative prediction errors are distributed in a wider range from 90.75% to 287.03%, with the average error magnitude equal to 53.30%. This regression error is mainly caused by the inadequacy of microstructure characterization for low-permeable rock samples, instead of under-fitting, over-fitting, or other training problems. Compared to the popular PNM [7, 113], which usually provides permeability evaluations with relative errors around ±40% for low-permeable rocks, the accuracy of the fitted neural network model is acceptable. On the other hand, this machine learning-based surrogate model also possesses an excellent generalization performance to predict permeability spanning four orders of magnitude for natural reservoir rocks.
Data-driven model II
In this part, the optimal feature set obtained in Sect. 3.4.4 is used as predictor variables to fit another data-driven model for permeability evaluation. This feature set contains nine morphological descriptors, as listed in Table 3. An FNN model is trained by minimizing its loss function, and its best training state can be reflected by the validation performance, as illustrated in Fig. 20a. Once the training process is completed, the FNN model is able to predict permeability for new observations, and the results are plotted in Fig. 20b. It seems that this regression model is comparable to data-driven model I in terms of prediction accuracy.
Fig. 20.
The data-driven FNN model II for permeability prediction
The relative prediction errors of data-driven model II are summarized in Table 3, and corresponding error distributions are plotted in Fig. 21. Compared to data-driven model I, data-driven model II possesses a slightly better prediction performance for low-permeable rock samples, but its prediction accuracy for high-permeable rock samples drops slightly. It may imply that blue descriptors are more accurate for microstructure characterization of high-permeable rocks, while green descriptors are more powerful to capture microstructural complexities of low-permeable rocks. Specifically, the relative prediction error varies from 50.66% to 46.79% for high-permeable rocks, with a average error magnitude equal to 8.80%. As to low-permeable rock samples, the range of relative evaluation error is from 85.82% to 272.81%, with the average error magnitude up to 49.73%. Although the prediction performances of data-driven model I and II are rather similar, the former (8 predictor variables) has much less predictor variables than the latter (83 predictor variables), and thus, data-driven model I can be considered to be a more cost-effective surrogate.
Fig. 21.
Relative error distributions of the permeability values evaluated from the data-driven model II
Generally, the pore network systems inside low-permeable rocks exhibit strong randomness, complexity, and heterogeneity, which makes it extremely difficult to achieve accurate and complete characterization of the internal microstructures. For high-permeable rock samples, the selected morphological descriptors can well represent the microstructural complexities for permeability evaluation. However, for low-permeable rock samples, more powerful morphological descriptors may be required to completely capture the microstructural characteristics related to permeability.
Comparisons
To examine the proposed data-driven framework in exploring the microstructure–permeability relation for porous media, the predictive models constructed through microstructural characterization and pore-scale simulation are compared with two popular empirical/analytical formulas in this section. Table 3 summarizes the performances of different predictive models in permeability evaluation. Generally, these two explicit formulas are significantly inferior to the data-driven models in terms of evaluation accuracy and generalization capacity.
Kozeny–Carman relation
The semi-empirical Kozeny–Carman relation [21, 25, 35] is one of the best-known formulas to estimate permeability, given by
| 15 |
where , S, and are the porosity, specific surface area, and tortuosity of a porous medium, respectively; and c is a dimensionless coefficient called Kozeny’s constant.
Kozeny’s constant c is an unknown coefficient, and its value can significantly vary with the microstructural characteristics of porous media. Without a universal value, c is usually estimated by empirically fitting numerical or experimental data for specific types of porous media [115]. For beds packed with spherical particles, c is around 2.50 [57]. Due to microstructural complexities, c should be larger than 2.50 for natural porous rocks. Besides, the three predictor variables , S, and involved in Kozeny–Carman equation are also contained in the feature selection results in Sect. 3.4.3 and Sect. 3.4.4. Here, the selected descriptors D2 (effective porosity), D3 (specific surface area), and D6 (geometrical tortuosity) are substituted into Eq. (15) to take the place of , S, and , respectively, and then, c is determined to be 4.26 by fitting the permeability results evaluated from lattice Boltzmann simulations.
As shown in Fig. 22a, Kozeny–Carman relation is unable to provide accurate predictions for a wide range of permeability values. Roughly, its prediction accuracy is acceptable when the permeability value is larger than 1,000 md, as illustrated by the data points in the green ellipse in Fig. 22a. The prediction error varies from 28.56% to 277.81%, with the average error magnitude up to 48.58%. However, Kozeny–Carman relation becomes much less reliable when it comes to lower permeable samples ( 1,000 md). Permeability values are systematically overestimated, as illustrated by the data points in the orange ellipse in Fig. 22a. The prediction error varies widely from 5.14% to 2393.70%, with the average error as high as 334.58%, which can be seen in Fig. 23b. In general, Kozeny–Carman relation possesses high level of uncertainty due to empirical selection of the adjustable coefficient c. Besides, the intrinsic microstructure–permeability mapping is also not fully represented by Kozeny–Carman relation, especially for low-permeable rocks, because only three simple morphological descriptors (namely, , S, and ) are used, which is far from sufficient to capture the microstructure complexities of natural porous rocks.
Fig. 22.
Permeability predictions obtained from (a) Kozeny–Carman relation, and (b) Berg’s relation (D12, D13, D14, D15, and D16 are the average pore size d results computed from different methods, and they are all used to replace in Eq. (16) for permeability evaluation)
Fig. 23.
Relative error distributions of the permeability values estimated from Kozeny–Carman relation
Berg’s relation
Inspired by Kozeny–Carman equation, Berg [10] derived a physical relation from the measurable microstructural descriptors of porous media, without introducing any tuning parameters or free constants. Berg’s relation reproduces Darcy’s law for idealized pipe flow, but it is also used to evaluate permeability for natural porous rocks, which is mathematically expressed as follows:
| 16 |
where is the effective porosity to describe the fractional volume conducting flow, denotes a characteristic length related to hydraulic pore radius, is the constriction factor to represent the fluctuation in local hydraulic pore radii, and denotes the tortuosity to quantify the effective length of streamlines.
Here, the selected morphological descriptors D2 (effective porosity), D6 (geometrical tortuosity), and D8 (constriction factor) are substituted into Eq. (16) to replace , , and respectively. Average pore size d (such as D12, D13, D14, D15, and D16) is used to approximate the characteristic length , and permeability results are then estimated from Berg’s relation, as shown in Fig. 22b. Five different methods are used to determine average pore sizes for the porous media samples used in this work, and the results from them are not consistent with each other, as illustrated in Fig. 5a. Obviously, the average pore size (D15) determined from the random point method is the most suitable estimation of for permeability evaluation via Berg’s relation, as can be seen from Fig. 22b. It should be emphasized that D15 is also one of the selected descriptors contained in the feature selection results in Sect. 3.4.4, which further substantiates the rationality of the feature selection results.
Here, the permeability results (the orange dots in Fig. 24) obtained by substituting D2, D6, D8, and D15 to into Eq. (16) are used to assess the performance of Berg’s relation. As shown in Fig. 24, the relative error distributions of the permeability values estimated from Berg’s relation are plotted. For high-permeable rock samples, Berg’s relation exhibits an acceptable accuracy, and the relative evaluation error varies from 66.04% to 93.66%, with the average error magnitude of 32.06%. However, the relative evaluation error increases significantly when it comes to low-permeable rock samples, as can be seen from Fig. 24b. The relative evaluation error varies from 79.10% to 1839.49%, with the average error magnitude up to 102.11%. Compared with Kozeny–Carman relation, Berg’s model exhibits a better prediction performance, and its greatest success is the exclusion of empirical parameter by introducing constriction factor . However, the inherent microstructure–permeability mapping is still not fully described by Berg’s relation, especially for low-permeable porous rocks.
Fig. 24.
Relative error distributions of the permeability values estimated from Berg’s relation
Discussion and conclusions
Discussion
As demonstrated in Sect. 3 and Sect. 4, it is an effective route to investigate the microstructure–property relationships through feature selection and data-driven modeling. Using the optimal feature set as the predictor variables for data-driven regression, a highly cost-effective model can be obtained with excellent prediction performance, and new insights into the microstructure–property linkage can also be gained from the feature selection results. However, achieving the above objectives is on the condition that the available feature pool can provide a comprehensive characterization of porous microstructures in an explicit expression, from which the optimal feature set that best represents the macroscopic physical property can be picked out through feature selection.
Generally, the data-driven surrogate models are of high accuracy and reliability to represent the microstructure–permeability mapping for high-permeable rocks. When it comes to low-permeable rocks, which usually possess more complicated internal pore network systems, neither the data-driven models nor explicit formulas can guarantee a high accuracy of permeability evaluation. The primary reason for such disparate performance is that the selected morphological descriptors capture the major microstructural characteristics that are important to permeability but also neglect some microstructural details, and such informatics loss can lead to increased uncertainty in permeability evolution when porous microstructures become more complicated.
Although there are many well-designed morphological descriptors (as listed in Table 1), extremely complicated microstructures could be beyond their capacity scope for accurate quantitative characterization. To cope with this inadequate characterization problem, effective descriptors should be specially developed from new perspectives, and microstructural details should also be preserved as much as possible by the new descriptors, which all put forward higher requests for quantitative microstructure analysis. Also, piecewise analysis can be adopted to establish microstructure–permeability mappings for different permeability ranges, because global predictive models are usually less accurate for low-permeable rock samples (as illustrated in Table 3) and the requirement of quantitative characterization also becomes higher as microstructural complexity increases.
Finally, the average computational costs of different predictive models to evaluate the permeability values of porous media samples used in this study are recorded in Table 4. It can be seen that the proposed data-driven models are able to provide instant predictions, which is times faster that the lattice Boltzmann simulation of pore-scale fluid flow.
Table 4.
The average computational costs of different predictive models for evaluating permeability of porous media samples used in this study
| Lattice Boltzmann simulation | Data-driven predictive models | Kozeny–Carman relation | Berg’s relation | |
|---|---|---|---|---|
| 30860 |
Conclusions
The main contribution of this work is to present a novel data-driven computational framework to fundamentally investigate the microstructure–property relationships of porous media through feature selection and data-driven regression. This framework can not only construct cost-effective surrogate models with high prediction accuracy and strong generalization capacity, but also provide new insights into the mechanisms of how microstructural characteristics determine microscopic behaviors.
This study especially focuses on the microstructure–permeability mapping of natural porous rocks. A large number of 3D digital microstructure samples with a wide porosity range are acquired from microscopy imaging and stochastic reconstruction. Pore-scale fluid flow passing through porous media is numerically simulated using high-fidelity lattice Boltzmann models, to provide reliable references of permeability values. A broad variety of morphological descriptors are collected from an extensive literature survey, and they compose the feature pool that quantitatively characterizes porous microstructures from global, local, geometrical, and topological perspectives. A performance-oriented feature selection is conducted to identify and pick out the microstructural characteristics that are significant to permeability. Based on the optimal feature sets, data-driven models are rapidly fitted to approximate the microstructure–permeability mapping, and these surrogate models can reliably predict permeability value spanning four orders of magnitudes, which are greatly superior to commonly used empirical/analytical formulas in terms of evaluation accuracy and generalization ability.
In addition to constructing cost-effective models, feature selection is also greatly beneficial to understanding the microstructure–permeability relation. By comparing the three categories of feature selection techniques (including filter, embedded, and wrapper methods), we found that the wrapper method is more applicable to exploring the microstructure–permeability linkage, because it is not only capable of identifying the joint importance of multiple features, but also effective for heterogeneous feature selection problems. According to the selected morphological descriptors, intrinsic permeability of porous media primarily depends on the microstructural characteristics in the following aspects: permeable pore volume, pore-solid interface, pore channel sinuosity, pore fractal dimension, pore coordination number, pore channel constriction, pore size, and percolation/connectivity degree. Besides, the proposed data-driven framework can be straightforwardly applied to analyze other physical properties (such as effective diffusivity, thermal conductivity, formation factor, and effective elastic moduli) of porous media by linking them to relevant microstructural informatics of importance.
Acknowledgements
The authors would like to acknowledge the support of EPSRC grant: PURIFY (EP/V000756/1), Swansea University (FSE Impact Fund), Higher Education Funding Council for Wales (COVID-19 Higher Education Student Support Fund), and Great Britain China Centre (Chinese Students Award).
Declarations
Conflict of interest
The authors declare that they have no conflict of interest in this paper.
Footnotes
Publisher's Note
Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
References
- 1.Adler PM. Porous media: geometry and transports. Boston: Butterworth-Heinemann; 1992. [Google Scholar]
- 2.Agrawal A, Choudhary A. Deep materials informatics: Applications of deep learning in materials science. MRS Commun. 2019;9(3):779–792. doi: 10.1557/mrc.2019.73. [DOI] [Google Scholar]
- 3.Andrä H, Combaret N, Dvorkin J, Glatt E, Han J, Kabel M, Keehm Y, Krzikalla F, Lee M, Madonna C, et al. Digital rock physics benchmarks-part ii: computing effective properties. Comput Geosci. 2013;50:33–43. doi: 10.1016/j.cageo.2012.09.008. [DOI] [Google Scholar]
- 4.Anovitz LM, Cole DR. Characterization and analysis of porosity and pore structures. Rev Mineral Geochem. 2015;80(1):61–164. doi: 10.2138/rmg.2015.80.04. [DOI] [Google Scholar]
- 5.Armstrong RT, McClure JE, Robins V, Liu Z, Arns CH, Schlüter S, Berg S. Porous media characterization using minkowski functionals: theories, applications and future directions. Transp Porous Media. 2019;130(1):305–335. doi: 10.1007/s11242-018-1201-4. [DOI] [Google Scholar]
- 6.Bache K, Lichman M (2013) UCI machine learning repository
- 7.Baychev TG, Jivkov AP, Rabbani A, Raeini AQ, Xiong Q, Lowe T, Withers PJ. Reliability of algorithms interpreting topological and geometric properties of porous media for pore network modelling. Transp Porous Media. 2019;128(1):271–301. doi: 10.1007/s11242-019-01244-8. [DOI] [Google Scholar]
- 8.Bear J (2013) Dynamics of fluids in porous media. Courier Corporation
- 9.Bentz DP, Martys NS. Hydraulic radius and transport in reconstructed model three-dimensional porous media. Transp Porous Media. 1994;17(3):221–238. doi: 10.1007/BF00613583. [DOI] [Google Scholar]
- 10.Berg CF. Permeability description by characteristic length, tortuosity, constriction and porosity. Transp Porous Media. 2014;103(3):381–400. doi: 10.1007/s11242-014-0307-6. [DOI] [Google Scholar]
- 11.Berryman JG, Blair SC. Kozeny-carman relations and image processing methods for estimating darcy’s constant. J Appl Phys. 1987;62(6):2221–2228. doi: 10.1063/1.339497. [DOI] [Google Scholar]
- 12.Bertei A, Nucci B, Nicolella C (2013) Effective transport properties in random packings of spheres and agglomerates
- 13.Bignonnet F. Efficient fft-based upscaling of the permeability of porous media discretized on uniform grids with estimation of rve size. Comput Methods Appl Mech Eng. 2020;369:113237. doi: 10.1016/j.cma.2020.113237. [DOI] [Google Scholar]
- 14.Biswal B, Manwart C, Hilfer R. Three-dimensional local porosity analysis of porous media. Phys A. 1998;255(3–4):221–241. doi: 10.1016/S0378-4371(98)00111-3. [DOI] [Google Scholar]
- 15.Blair SC, Berge PA, Berryman JG. Using two-point correlation functions to characterize microgeometry and estimate permeabilities of sandstones and porous glass. J Geophys Res Solid Earth. 1996;101(B9):20359–20375. doi: 10.1029/96JB00879. [DOI] [Google Scholar]
- 16.Blunt MJ. Multiphase flow in permeable media: A pore-scale perspective. Cambridge: Cambridge University Press; 2017. [Google Scholar]
- 17.Blunt MJ, Bijeljic B, Dong H, Gharbi O, Iglauer S, Mostaghimi P, Paluszny A, Pentland C. Pore-scale imaging and modelling. Adv Water Resour. 2013;51:197–216. doi: 10.1016/j.advwatres.2012.03.003. [DOI] [Google Scholar]
- 18.Bostanabad R, Zhang Y, Li X, Kearney T, Brinson LC, Apley DW, Liu WK, Chen W. Computational microstructure characterization and reconstruction: Review of the state-of-the-art techniques. Prog Mater Sci. 2018;95:1–41. doi: 10.1016/j.pmatsci.2018.01.005. [DOI] [Google Scholar]
- 19.Breiman L. Random forests. Mach Learn. 2001;45(1):5–32. doi: 10.1023/A:1010933404324. [DOI] [Google Scholar]
- 20.Bultreys T (2016) Savonnières carbonate. http://www.digitalrocksportal.org/projects/72
- 21.Carman PC. Fluid flow through granular beds. Trans Inst Chem Eng. 1937;50:150–166. [Google Scholar]
- 22.Cecen A, Dai H, Yabansu YC, Kalidindi SR, Song L. Material structure-property linkages using three-dimensional convolutional neural networks. Acta Mater. 2018;146:76–84. doi: 10.1016/j.actamat.2017.11.053. [DOI] [Google Scholar]
- 23.Cecen A, Wargo E, Hanna A, Turner D, Kalidindi S, Kumbur E. 3-d microstructure analysis of fuel cell materials: spatial distributions of tortuosity, void size and diffusivity. J Electrochem Soc. 2012;159(3):B299. doi: 10.1149/2.068203jes. [DOI] [Google Scholar]
- 24.Chen H, Chen S, Matthaeus WH. Recovery of the navier-stokes equations using a lattice-gas boltzmann method. Phys Rev A. 1992;45(8):R5339. doi: 10.1103/PhysRevA.45.R5339. [DOI] [PubMed] [Google Scholar]
- 25.Clennell MB. Tortuosity: a guide through the maze. Geol Soc Lond Spec Publ. 1997;122(1):299–344. doi: 10.1144/GSL.SP.1997.122.01.18. [DOI] [Google Scholar]
- 26.Coker DA, Torquato S, Dunsmuir JH. Morphology and physical properties of fontainebleau sandstone via a tomographic analysis. J Geophys Res Solid Earth. 1996;101(B8):17497–17506. doi: 10.1029/96JB00811. [DOI] [Google Scholar]
- 27.Cosenza P, Prêt D, Zamora M. Effect of the local clay distribution on the effective electrical conductivity of clay rocks. J Geophys Res Solid Earth. 2015;120(1):145–168. doi: 10.1002/2014JB011429. [DOI] [Google Scholar]
- 28.Costa A. Permeability-porosity relationship: a reexamination of the kozeny-carman equation based on a fractal pore-space geometry assumption. Geophys Res Lett. 2006;33(2):5. doi: 10.1029/2005GL025134. [DOI] [Google Scholar]
- 29.Cui S, Fu J, Cen S, Thomas HR, Li C. The correlation between statistical descriptors of heterogeneous materials. Comput Methods Appl Mech Eng. 2021;384:113948. doi: 10.1016/j.cma.2021.113948. [DOI] [Google Scholar]
- 30.Delerue J, Perrier E, Yu Z, Velde B. New algorithms in 3d image analysis and their application to the measurement of a spatialized pore size distribution in soils. Phys Chem Earth Part A. 1999;24(7):639–644. doi: 10.1016/S1464-1895(99)00093-9. [DOI] [Google Scholar]
- 31.Dong H, Gao P, Ye G. Characterization and comparison of capillary pore structures of digital cement pastes. Mater Struct. 2017;50(2):154. doi: 10.1617/s11527-017-1023-9. [DOI] [Google Scholar]
- 32.Erofeev A, Orlov D, Ryzhov A, Koroteev D. Prediction of porosity and permeability alteration based on machine learning algorithms. Transp Porous Media. 2019;128(2):677–700. doi: 10.1007/s11242-019-01265-3. [DOI] [Google Scholar]
- 33.Fu J, Cui S, Cen S, Li C. Statistical characterization and reconstruction of heterogeneous microstructures using deep neural network. Comput Methods Appl Mech Eng. 2021;373:113516. doi: 10.1016/j.cma.2020.113516. [DOI] [Google Scholar]
- 34.Fu J, Dong J, Wang Y, Ju Y, Owen DRJ, Li C. Resolution effect: An error correction model for intrinsic permeability of porous media estimated from lattice boltzmann method. Transp Porous Media. 2020;132(3):627–656. doi: 10.1007/s11242-020-01406-z. [DOI] [Google Scholar]
- 35.Fu J, Thomas HR, Li C. Tortuosity of porous media: Image analysis and physical simulation. Earth-Sci Rev. 2020;2:103439. [Google Scholar]
- 36.Fu J, Wang M, Xiao D, Zhong S, Ge X, Ben E (2023) Hierarchical reconstruction of 3d well-connected porous media from 2d exemplars using statistics-informed neural network. Comput Methods Appl Mech Eng
- 37.Fu J, Xiao D, Fu R, Li C, Zhu C, Arcucci R, Navon IM. Physics-data combined machine learning for parametric reduced-order modelling of nonlinear dynamical systems in small-data regimes. Comput Methods Appl Mech Eng. 2023;404:115771. doi: 10.1016/j.cma.2022.115771. [DOI] [Google Scholar]
- 38.Fu J, Xiao D, Li D, Thomas HR, Li C. Stochastic reconstruction of 3d microstructures from 2d cross-sectional images using machine learning-based characterization. Comput Methods Appl Mech Eng. 2022;390:114532. doi: 10.1016/j.cma.2021.114532. [DOI] [Google Scholar]
- 39.Géraud Y. Variations of connected porosity and inferred permeability in a thermally cracked granite. Geophys Res Lett. 1994;21(11):979–982. doi: 10.1029/94GL00642. [DOI] [Google Scholar]
- 40.Goodarzi M, Dejaegher B, Heyden YV. Feature selection methods in qsar studies. J AOAC Int. 2012;95(3):636–651. doi: 10.5740/jaoacint.SGE_Goodarzi. [DOI] [PubMed] [Google Scholar]
- 41.Guest JK, Prévost JH. Design of maximum permeability material structures. Comput Methods Appl Mech Eng. 2007;196(4–6):1006–1017. doi: 10.1016/j.cma.2006.08.006. [DOI] [Google Scholar]
- 42.Guyon I, Elisseeff A. An introduction to variable and feature selection. J Mach Learn Res. 2003;3:1157–1182. [Google Scholar]
- 43.Guyon I, Gunn S, Nikravesh M, Zadeh LA. Feature extraction: foundations and applications, Berlin: Springer; 2008. [Google Scholar]
- 44.Herring A, Sheppard A, Turner M, Beeching L (2018) Multiphase flows in sandstones. http://www.digitalrocksportal.org/projects/135
- 45.Hilfer R. Review on scale dependent characterization of the microstructure of porous media. Transp Porous Media. 2002;46(2–3):373–390. doi: 10.1023/A:1015014302642. [DOI] [Google Scholar]
- 46.Holzer L, Iwanschitz B, Hocker T, Münch B, Prestat M, Wiedenmann D, Vogt U, Holtappels P, Sfeir J, Mai A, et al. Microstructure degradation of cermet anodes for solid oxide fuel cells: Quantification of nickel grain growth in dry and in humid atmospheres. J Power Sources. 2011;196(3):1279–1294. doi: 10.1016/j.jpowsour.2010.08.017. [DOI] [Google Scholar]
- 47.Holzer L, Wiedenmann D, Münch B, Keller L, Prestat M, Gasser P, Robertson I, Grobéty B. The influence of constrictivity on the effective transport properties of porous layers in electrolysis and fuel cells. J Mater Sci. 2013;48(7):2934–2952. doi: 10.1007/s10853-012-6968-z. [DOI] [Google Scholar]
- 48.Hong J, Liu J. Rapid estimation of permeability from digital rock using 3d convolutional neural network. Comput Geosci. 2020;24:1523–1539. doi: 10.1007/s10596-020-09941-w. [DOI] [Google Scholar]
- 49.Hormann K, Baranau V, Hlushkou D, Höltzel A, Tallarek U. Topological analysis of non-granular, disordered porous media: determination of pore connectivity, pore coordination, and geometric tortuosity in physically reconstructed silica monoliths. New J Chem. 2016;40(5):4187–4199. doi: 10.1039/C5NJ02814K. [DOI] [Google Scholar]
- 50.Iassonov P, Gebrenegus T, Tuller M. Segmentation of x-ray computed tomography images of porous materials: A crucial step for characterization and quantitative analysis of pore structures. Water Resour Res. 2009;45(9):6. doi: 10.1029/2009WR008087. [DOI] [Google Scholar]
- 51.ImageJ (2016) Website. https://imagej.net/Welcome
- 52.Ioannidis M, Kwiecien M, Chatzis I. Statistical analysis of the porous microstructure as a method for estimating reservoir permeability. J Petrol Sci Eng. 1996;16(4):251–261. doi: 10.1016/S0920-4105(96)00044-7. [DOI] [Google Scholar]
- 53.Jiao Y, Stillinger F, Torquato S. A superior descriptor of random textures and its predictive capacity. Proc Natl Acad Sci. 2009;106(42):17634–17639. doi: 10.1073/pnas.0905919106. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 54.Jin G, Patzek T, Silin D (2004) Direct prediction of the absolute permeability of unconsolidated and consolidated reservoir rock. spe 90084. In (2003) SPE Annual Technical Conference and Exhibition (Houston. Texas, USA), SPE
- 55.Kamrava S, Tahmasebi P, Sahimi M. Linking morphology of porous media to their macroscopic permeability by deep learning. Transp Porous Media. 2020;131(2):427–448. doi: 10.1007/s11242-019-01352-5. [DOI] [Google Scholar]
- 56.Karimpouli S, Tahmasebi P. Image-based velocity estimation of rock using convolutional neural networks. Neural Netw. 2019;111:89–97. doi: 10.1016/j.neunet.2018.12.006. [DOI] [PubMed] [Google Scholar]
- 57.Kaviany M. Principles of heat transfer in porous media. Berlin: Springer Science & Business Media; 2012. [Google Scholar]
- 58.Knudby C, Carrera J. On the relationship between indicators of geostatistical, flow and transport connectivity. Adv Water Resour. 2005;28(4):405–421. doi: 10.1016/j.advwatres.2004.09.001. [DOI] [Google Scholar]
- 59.Kohanpur AH, Valocchi A, Crandall D (2019) Micro-ct images of a heterogeneous mt. simon sandstone sample. http://www.digitalrocksportal.org/projects/247
- 60.Kohavi R, John GH, et al. Wrappers for feature subset selection. Artif Intell. 1997;97(1–2):273–324. doi: 10.1016/S0004-3702(97)00043-X. [DOI] [Google Scholar]
- 61.Koponen A, Kataja M, Timonen J. Permeability and effective porosity of porous media. Phys Rev E. 1997;56(3):3319. doi: 10.1103/PhysRevE.56.3319. [DOI] [Google Scholar]
- 62.Krüger T, Kusumaatmaja H, Kuzmin A, Shardt O, Silva G, Viggen EM. The lattice boltzmann method: principles and practice. Berlin: Springer; 2016. [Google Scholar]
- 63.Kuhn M, Johnson K, et al. Applied predictive modeling. Berlin: Springer; 2013. [Google Scholar]
- 64.Kutay ME, Aydilek AH, Masad E. Laboratory validation of lattice boltzmann method for modeling pore-scale flow in granular materials. Comput Geotech. 2006;33(8):381–395. doi: 10.1016/j.compgeo.2006.08.002. [DOI] [Google Scholar]
- 65.Latief F, Biswal B, Fauzi U, Hilfer R. Continuum reconstruction of the pore scale microstructure for fontainebleau sandstone. Phys A. 2010;389(8):1607–1618. doi: 10.1016/j.physa.2009.12.006. [DOI] [Google Scholar]
- 66.Lehmann P, Berchtold M, Ahrenholz B, Tölke J, Kaestner A, Krafczyk M, Flühler H, Künsch H. Impact of geometrical properties on permeability and fluid phase distribution in porous media. Adv Water Resour. 2008;31(9):1188–1204. doi: 10.1016/j.advwatres.2008.01.019. [DOI] [Google Scholar]
- 67.Li J, Cheng K, Wang S, Morstatter F, Trevino RP, Tang J, Liu H. Feature selection: a data perspective. ACM Comput Surv (CSUR) 2017;50(6):1–45. doi: 10.1145/3136625. [DOI] [Google Scholar]
- 68.Li X, Liu Z, Cui S, Luo C, Li C, Zhuang Z. Predicting the effective mechanical property of heterogeneous materials by image based modeling and deep learning. Comput Methods Appl Mech Eng. 2019;347:735–753. doi: 10.1016/j.cma.2019.01.005. [DOI] [Google Scholar]
- 69.Liang Z, Ioannidis M, Chatzis I. Permeability and electrical conductivity of porous media from 3d stochastic replicas of the microstructure. Chem Eng Sci. 2000;55(22):5247–5262. doi: 10.1016/S0009-2509(00)00142-1. [DOI] [Google Scholar]
- 70.Lindquist WB, Venkatarangan A, Dunsmuir J, Wong T-F. Pore and throat size distributions measured from synchrotron x-ray tomographic images of fontainebleau sandstones. J Geophys Res Solid Earth. 2000;105(B9):21509–21527. doi: 10.1029/2000JB900208. [DOI] [Google Scholar]
- 71.Liu J, Pereira GG, Liu Q, Regenauer-Lieb K. Computational challenges in the analyses of petrophysics using microtomography and upscaling: a review. Comput Geosci. 2016;89:107–117. doi: 10.1016/j.cageo.2016.01.014. [DOI] [Google Scholar]
- 72.Loh W-Y. Regression tress with unbiased variable selection and interaction detection. Stat Sin. 2002;2:361–386. [Google Scholar]
- 73.Luhmann AJ, Tutolo BM, Bagley BC, Mildner DF, Seyfried WE, Jr, Saar MO. Permeability, porosity, and mineral surface area changes in basalt cores induced by reactive transport of co 2-rich brine. Water Resour Res. 2017;53(3):1908–1927. doi: 10.1002/2016WR019216. [DOI] [Google Scholar]
- 74.Łydżba D, Różański A, Sevostianov I, Stefaniuk D. A new methodology for evaluation of thermal or electrical conductivity of the skeleton of a porous material. Int J Eng Sci. 2021;158:103397. doi: 10.1016/j.ijengsci.2020.103397. [DOI] [Google Scholar]
- 75.Moctezuma-Berthier A, Vizika O, Adler PM. Macroscopic conductivity of vugular porous media. Transp Porous Media. 2002;49(3):313–332. doi: 10.1023/A:1016297220013. [DOI] [Google Scholar]
- 76.Moré JJ. Numerical analysis. Berlin: Springer; 1978. The levenberg-marquardt algorithm: implementation and theory; pp. 105–116. [Google Scholar]
- 77.Muche L, Stoyan D. Contact and chord length distributions of the poisson voronoi tessellation. J Appl Probab. 1992;2:467–471. doi: 10.2307/3214584. [DOI] [Google Scholar]
- 78.Münch B, Holzer L. Contradicting geometrical concepts in pore size analysis attained with electron microscopy and mercury intrusion. J Am Ceram Soc. 2008;91(12):4059–4067. doi: 10.1111/j.1551-2916.2008.02736.x. [DOI] [Google Scholar]
- 79.N’Diaye M, Degeratu C, Bouler J-M, Chappard D. Biomaterial porosity determined by fractal dimensions, succolarity and lacunarity on microcomputed tomographic images. Mater Sci Eng, C. 2013;33(4):2025–2030. doi: 10.1016/j.msec.2013.01.020. [DOI] [PubMed] [Google Scholar]
- 80.Nordlund M, Penha DJL, Stolz S, Kuczaj A, Winkelmann C, Geurts BJ. A new analytical model for the permeability of anisotropic structured porous media. Int J Eng Sci. 2013;68:38–60. doi: 10.1016/j.ijengsci.2013.01.003. [DOI] [Google Scholar]
- 81.Okabe H, Blunt MJ. Prediction of permeability for porous media reconstructed using multiple-point statistics. Phys Rev E. 2004;70(6):066135. doi: 10.1103/PhysRevE.70.066135. [DOI] [PubMed] [Google Scholar]
- 82.Paterson M. The equivalent channel model for permeability and resistivity in fluid-saturated rock-a re-appraisal. Mech Mater. 1983;2(4):345–352. doi: 10.1016/0167-6636(83)90025-X. [DOI] [Google Scholar]
- 83.Pia G, Sanna U. An intermingled fractal units model and method to predict permeability in porous rock. Int J Eng Sci. 2014;75:31–39. doi: 10.1016/j.ijengsci.2013.11.002. [DOI] [Google Scholar]
- 84.Quintanilla J, Torquato S. Local volume fraction fluctuations in random media. J Chem Phys. 1997;106(7):2741–2751. doi: 10.1063/1.473414. [DOI] [Google Scholar]
- 85.Rabbani A, Babaei M, Shams R, Da Wang Y, Chung T. Deepore: a deep learning workflow for rapid and comprehensive characterization of porous materials. Adv Water Resour. 2020;2:103787. doi: 10.1016/j.advwatres.2020.103787. [DOI] [Google Scholar]
- 86.Robnik-Šikonja M, Kononenko I. Theoretical and empirical analysis of relieff and rrelieff. Mach Learn. 2003;53(1–2):23–69. doi: 10.1023/A:1025667309714. [DOI] [Google Scholar]
- 87.Röding M, Ma Z, Torquato S. Predicting permeability via statistical learning on higher-order microstructural information. Sci Rep. 2020;10(1):1–17. doi: 10.1038/s41598-020-72085-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 88.Rubinstein J, Torquato S. Diffusion-controlled reactions: Mathematical formulation, variational principles, and rigorous bounds. J Chem Phys. 1988;88(10):6372–6380. doi: 10.1063/1.454474. [DOI] [Google Scholar]
- 89.Rubinstein J, Torquato S. Flow in random porous media: mathematical formulation, variational principles, and rigorous bounds. J Fluid Mech. 1989;206:25–46. doi: 10.1017/S0022112089002211. [DOI] [Google Scholar]
- 90.Santos JE, Xu D, Jo H, Landry CJ, Prodanović M, Pyrcz MJ. Poreflow-net: A 3d convolutional neural network to predict fluid flow through porous media. Adv Water Resour. 2020;138:103539. doi: 10.1016/j.advwatres.2020.103539. [DOI] [Google Scholar]
- 91.Saxena N, Hofmann R, Alpak FO, Berg S, Dietderich J, Agarwal U, Tandon K, Hunter S, Freeman J, Wilson OB. References and benchmarks for pore-scale flow simulated using micro-ct images of porous media and digital rocks. Adv Water Resour. 2017;109:211–235. doi: 10.1016/j.advwatres.2017.09.007. [DOI] [Google Scholar]
- 92.Scanziani A, Singh K, Blunt M (2018) Water-wet three-phase flow micro-ct tomograms. http://www.digitalrocksportal.org/projects/167
- 93.Schlüter S, Sheppard A, Brown K, Wildenschild D. Image processing of multiphase images obtained via x-ray microtomography: a review. Water Resour Res. 2014;50(4):3615–3639. doi: 10.1002/2014WR015256. [DOI] [Google Scholar]
- 94.Schulz E, Speekenbrink M, Krause A. A tutorial on gaussian process regression: modelling, exploring, and exploiting functions. J Math Psychol. 2018;85:1–16. doi: 10.1016/j.jmp.2018.03.001. [DOI] [Google Scholar]
- 95.Sevostianova E, Leinauer B, Sevostianov I. Quantitative characterization of the microstructure of a porous material in the context of tortuosity. Int J Eng Sci. 2010;48(12):1693–1701. doi: 10.1016/j.ijengsci.2010.06.017. [DOI] [Google Scholar]
- 96.Srisutthiyakorn N (2016) Deep-learning methods for predicting permeability from 2d/3d binary-segmented images. In: SEG technical program expanded abstracts 2016, pages 3042–3046. Society of Exploration Geophysicists
- 97.Sudakov O, Burnaev E, Koroteev D. Driving digital rock towards machine learning: Predicting permeability with gradient boosting and deep neural networks. Comput Geosci. 2019;127:91–98. doi: 10.1016/j.cageo.2019.02.002. [DOI] [Google Scholar]
- 98.Sukop MC, Huang H, Alvarez PF, Variano EA, Cunningham KJ. Evaluation of permeability and non-darcy flow in vuggy macroporous limestone aquifer samples with lattice boltzmann methods. Water Resour Res. 2013;49(1):216–230. doi: 10.1029/2011WR011788. [DOI] [Google Scholar]
- 99.Tembely M, AlSumaiti AM, Alameri W. A deep learning perspective on predicting permeability in porous media from network modeling to direct simulation. Comput Geosci. 2020;24:1541–1556. doi: 10.1007/s10596-020-09963-4. [DOI] [Google Scholar]
- 100.Tian J, Qi C, Sun Y, Yaseen ZM. Surrogate permeability modelling of low-permeable rocks using convolutional neural networks. Comput Methods Appl Mech Eng. 2020;366:103–113. doi: 10.1016/j.cma.2020.113103. [DOI] [Google Scholar]
- 101.Tian J, Qi C, Sun Y, Yaseen ZM, Pham BT. Permeability prediction of porous media using a combination of computational fluid dynamics and hybrid machine learning methods. Eng Comput. 2020;2:1–17. [Google Scholar]
- 102.Torquato S. Statistical description of microstructures. Annu Rev Mater Res. 2002;32(1):77–111. doi: 10.1146/annurev.matsci.32.110101.155324. [DOI] [Google Scholar]
- 103.Torquato S. Random heterogeneous materials: microstructure and macroscopic properties. Berlin: Springer Science & Business Media; 2013. [Google Scholar]
- 104.van der Linden JH, Narsilio GA, Tordesillas A. Machine learning framework for analysis of transport through complex networks in porous, granular media: a focus on permeability. Phys Rev E. 2016;94(2):022904. doi: 10.1103/PhysRevE.94.022904. [DOI] [PubMed] [Google Scholar]
- 105.Vogel H-J, Weller U, Schlüter S. Quantification of soil structure based on minkowski functions. Comput Geosci. 2010;36(10):1236–1245. doi: 10.1016/j.cageo.2010.03.007. [DOI] [Google Scholar]
- 106.Wang J, Li Z, Yan S, Yu X, Ma Y, Ma L. Modifying the microstructure of algae-based active carbon and modelling supercapacitors using artificial neural networks. RSC Adv. 2019;9(26):14797–14808. doi: 10.1039/C9RA01255A. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 107.Wang M, Pan N. Modeling and prediction of the effective thermal conductivity of random open-cell porous foams. Int J Heat Mass Transf. 2008;51(5–6):1325–1331. doi: 10.1016/j.ijheatmasstransfer.2007.11.031. [DOI] [Google Scholar]
- 108.Wei H, Zhao S, Rong Q, Bao H. Predicting the effective thermal conductivities of composite materials and porous media by machine learning methods. Int J Heat Mass Transf. 2018;127:908–916. doi: 10.1016/j.ijheatmasstransfer.2018.08.082. [DOI] [Google Scholar]
- 109.Wu H, Fang W-Z, Kang Q, Tao W-Q, Qiao R. Predicting effective diffusivity of porous media from images by deep learning. Sci Rep. 2019;9(1):1–12. doi: 10.1038/s41598-019-56309-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 110.Wu J, Yin X, Xiao H. Seeing permeability from images: fast prediction with convolutional neural networks. Sci Bull. 2018;63(18):1215–1222. doi: 10.1016/j.scib.2018.08.006. [DOI] [PubMed] [Google Scholar]
- 111.Xia M, Fu J, Feng Y, Gong F, Jin Y. A particle-resolved heat-particle-fluid coupling model by dem-imb-lbm. J Rock Mech Geotech Eng. 2023;2:2. [Google Scholar]
- 112.Xia Y, Cai J, Perfect E, Wei W, Zhang Q, Meng Q. Fractal dimension, lacunarity and succolarity analyses on ct images of reservoir rocks for permeability prediction. J Hydrol. 2019;579:124198. doi: 10.1016/j.jhydrol.2019.124198. [DOI] [Google Scholar]
- 113.Xiong Q, Baychev TG, Jivkov AP. Review of pore network modelling of porous media: experimental characterisations, network constructions and applications to reactive transport. J Contam Hydrol. 2016;192:101–117. doi: 10.1016/j.jconhyd.2016.07.002. [DOI] [PubMed] [Google Scholar]
- 114.Xu H, Dikin DA, Burkhart C, Chen W. Descriptor-based methodology for statistical characterization and 3d reconstruction of microstructural materials. Comput Mater Sci. 2014;85:206–216. doi: 10.1016/j.commatsci.2013.12.046. [DOI] [Google Scholar]
- 115.Xu P, Yu B. Developing a new form of permeability and Kozeny–Carman constant for homogeneous porous media by means of fractal geometry. Adv Water Resour. 2008;31(1):74–81. doi: 10.1016/j.advwatres.2007.06.003. [DOI] [Google Scholar]
- 116.Yang W, Wang K, Zuo W. Neighborhood component feature selection for high-dimensional data. JCP. 2012;7(1):161–168. [Google Scholar]
- 117.Yegnanarayana B. Artificial neural networks. Delhi: PHI Learning Pvt Ltd.; 2009. [Google Scholar]
- 118.Yu B, Cheng P. A fractal permeability model for bi-dispersed porous media. Int J Heat Mass Transf. 2002;45(14):2983–2993. doi: 10.1016/S0017-9310(02)00014-5. [DOI] [Google Scholar]
- 119.Zeng Z, Fu J, Feng Y, Wang M. Revisiting the empirical particle-fluid coupling model used in dem-cfd by high-resolution dem-lbm-imb simulations: a 2d perspective. Int J Numer Anal Methods Geomech. 2023;2:2. [Google Scholar]






















