Skip to main content
Toxicological Research logoLink to Toxicological Research
. 2025 Nov 24;42(2):235–250. doi: 10.1007/s43188-025-00326-8

Artificial intelligence‑based quantitative analysis of hepatic fibrosis in carbon tetrachloride-induced mouse model of metabolic dysfunction-associated steatohepatitis

Jin-Hee Lee 1,3, Myung-Hwa Yang 1, Gyeongjin Han 1, Won Hoon Jung 2, Myung Ae Bae 2, Ji-Seok Han 1, Tae-Sung Koo 3, Jae-Woo Cho 1,✉
PMCID: PMC12945867  PMID: 41769373

Abstract

Liver fibrosis, a major histopathological indicator of chronic liver injury, is also a key feature of metabolic dysfunction-associated steatohepatitis. Its quantitative assessment in preclinical toxicology is frequently inconsistent and subjective. This study aimed to develop and validate multi-scale, patch-based convolutional neural network classification algorithms for automated fibrosis quantification in a carbon tetrachloride (CCl4)-induced mouse model. We sought to determine the optimal patch size for accurate predictions. Accordingly, male C57BL/6 mice (n = 19) were categorized into the following three groups: vehicle control (n = 5), high-fat diet (HFD) and CCl4 positive control (n = 9), and HFD and CCl4 with elafibranor (ELA) treatment (n = 5). Liver tissues were stained with Sirius-red, digitized as whole slide images, and cropped into patches of 32 × 32, 64 × 64, or 128 × 128 pixels. Each algorithm was trained, validated, and tested in an 8:1:1 ratio over 40 epochs with a batch size of 32 to classify fibrotic, normal, and background regions. All models performed robustly, with validation accuracies exceeding 98% and F1-scores above 0.96. Particularly, the 32 × 32 model exhibited the highest correlation with pathologist’s measurements (Spearman’s r = 0.9609; p < 0.05) and the most accurate estimation of absolute fibrotic area compared to expert assessments. This model also accurately detected the antifibrotic effects of ELA. These findings establish that the 32 × 32 patch-based classification approach provides a rapid, reproducible, and objective method for liver fibrosis quantification in preclinical toxicology, with strong potential for integration into digital pathology workflows.

Graphical abstract

graphic file with name 43188_2025_326_Figa_HTML.jpg

Supplementary Information

The online version contains supplementary material available at 10.1007/s43188-025-00326-8.

Keywords: Hepatic fibrosis, Metabolic dysfunction-associated steatohepatitis, Automated quantification, Deep learning, Digital pathology

Introduction

The prevalence of metabolic syndrome-related diseases, commonly known as lifestyle diseases, is continually increasing. Of these, metabolic dysfunction-associated steatotic liver disease (MASLD) and steatohepatitis (MASH) are major causes of liver fibrosis [1–3]. If these diseases worsen, they cause necrotic inflammation. Consequently, persistent liver damage triggers a chronic inflammatory response, ultimately leading to liver fibrosis [4]. Currently, no drugs for MASLD and MASH are approved by the Food and Drug Administration, and fibrosis is considered a disease that requires urgent therapeutic development [5].

Toxicity study is a crucial process in drug evaluation and development [6], with histopathological examination of experimental animals serving as a key process for diagnostic assessment [7]. Histopathological examination characterizes lesions with various staining techniques using tissue slides. Additionally, it analyzes pathological information and clinical correlations. This process involves assessing objective and quantitative pathological information, including cell count, lesion size, and distribution area. However, the reliability of this evaluation may vary depending on the pathologist’s prior knowledge, expertise, fatigue, and experience. Moreover, considering that hundreds or thousands of tissue slides need to be swiftly analyzed, standardized large-scale datasets that reduce variability from assessment by different pathologists are needed. However, such datasets are currently lacking [8].

To address these issues, digital pathology approaches are being introduced along with advancing medical imaging technology [9]. Furthermore, whole slide images (WSIs) are currently obtained as virtual microscope images due to the development of digital slide scanners. This enables pathologists to conduct in-depth tissue analyses via computer systems and provides a faster and more quantitative approach to analysis [10, 11]. Additionally, these digital images are used as a platform for applying artificial intelligence (AI) [12]. Notably, the rapid advancements in machine and deep learning algorithms have enabled lesion detection through images and automated classification, thereby significantly improving diagnostic speed and accuracy compared with traditional methods [13]. Moreover, deep learning technology has been widely used in the medical field, particularly since the development of the convolutional neural network (CNN). Unlike the artificial neural network that recognizes an image as a single piece of data and performs sub-optimally when the image is distorted, CNNs can divide an image into small filters and automatically extract features, thereby enabling pattern recognition even when the image is distorted or its location changes [14]. These AI-based image analysis techniques are widely used in clinical fields to accurately diagnose tumors and various cancer tissues [15, 16]. In non-clinical fields, they are used to develop classification models that detect the presence or absence of proliferative carcinogenic lesions in mice [17].

Over the past decade, clinical studies were consistently conducted to quantify liver fibrosis using AI [18], and the development of related rodent models was actively pursued. One study quantitatively assessed the nonlinear elastic collagen proportional area from non-invasive microscopic images using a CNN algorithm with a patch size of 512 × 512, demonstrating potential as a technological tool for deep learning with a model prediction accuracy of 84% [19]. Another study developed an automatic scoring algorithm model based on fibrosis severity. In this study, CNN classification algorithms with a patch size of 512 × 512 were used to automatically score pulmonary fibrosis and inflammation in hematoxylin and eosin-stained mouse tissues, achieving learning accuracies of 79.5% and 80.0% for fibrosis and inflammation, respectively [20]. Subsequently, the same research team used a CNN classification algorithm to automatically score liver fibrosis findings in rodent models stained with Masson’s trichrome, achieving a quantification accuracy of 86.3% with a patch size of 299 × 299 [21]. However, among these studies, none have identified the optimal patch size for quantifying liver fibrosis in mouse models by training algorithms with multiple patch images of smaller sizes [22].

Therefore, in this study, we aimed to automatically quantify liver fibrosis lesions in WSIs from MASH-induced mouse models stained with Sirius-red using CNN classification model algorithms.

Material and methods

This study aimed to develop algorithms using three patch-based approaches to classify liver fibrosis. Accordingly, we investigated normal liver tissue, fibrotic liver tissue induced by MASH, and partially recovered fibrotic liver tissue after treatment with elafibranor (ELA), which is an antifibrotic candidate. Each liver tissue was stained using the Sirius Red staining technique and converted into digital slides. The acquired WSIs were segmented into patches of sizes 32 × 32, 64 × 64, and 128 × 128, which were used to train the CNN algorithms. The dataset was randomly divided into training, validation, and test sets using an 8:1:1 ratio. Finally, the three developed algorithms were compared with the fibrotic regions on WSIs directly annotated by a pathologist to analyze their similarity.

Animal treatment and histological analysis

All animals were housed under controlled conditions with a 12-h light:12-h dark cycle, temperature (23 ± 1 °C), ventilation (10–12 air changes per hour), and humidity (55 ± 5%). Mice were kept in groups of 4–5 per cage with ad libitum access to food and water, following National Institutes of Health guidelines. C57BL/6 mice were purchased from Orient Bio (Seongnam, Republic of Korea) and housed in a barrier system at the Korea Research Institute of Chemical Technology (Daejeon, Republic of Korea). Cages were changed two times weekly, and all equipment and bedding were sterilized with an autoclave. Animals had free access to ultraviolet-sterilized filtered tap water. All animals were monitored daily for signs of morbidity or mortality throughout the experiment in accordance with the approved protocol. No moribund animals or unexpected deaths were recorded in any group, and all animals survived until the planned endpoint. All animals were euthanized as scheduled at the completion of the study. The Institutional Animal Care and Use Committee of the Korea Research Institute of Chemical Technology approved this experimental protocol (Permit Number: 2020-7A-02-02).

Male C57BL/6 mice, introduced at 6 weeks of age, were stabilized for 7 days before the start of the experiment. During the experimental period, they were categorized into the following three groups: the vehicle control (Purina, Gunpo, Republic of Korea—Fat: 4.5%, Protein: 20.12%, Fiber: 3.5%), positive control (high-fat diet [HFD] + carbon tetrachloride [CCl4]), and HFD + CCl4 + ELA groups. To evaluate the antifibrotic effect [23], animals were distributed as follows: vehicle control (n = 5), HFD + CCl4 group (n = 9), and HFD + CCl4 + ELA group (n = 5). The HFD used in the experiment contained 60 kcal% fat and was purchased from Research Diets, Inc. (New Brunswick, USA) under catalog number D12492. All animals were fed ad libitum. CCl4 (2% v/v in corn oil; Sigma–Aldrich, Cat#C8267) [24] was administered intraperitoneally two times weekly for 4 weeks. The initial injection was administered at 2 mL/kg, followed by 3 mL/kg for subsequent injections. The HFD + CCl4 + ELA group received ELA (MedKoo Biosciences, Cat#561377) for 14 consecutive days after fibrosis induction. ELA was mixed with 0.5% carboxymethylcellulose sodium solution, which served as the vehicle; the compound was ground in a glass mortar and sonicated to improve solubility, then administered orally at 20 mg/kg/day [25, 26]. The details of the animal treatment process are shown in Fig. 1.

Fig. 1.

Fig. 1

Overall process of the animal experiments. Acquisition, administration, and liver tissue collection in C57BL/c mice. CMC carboxyl methyl cellulose, CCl4 carbon tetrachloride, ELA elafibranor

Two weeks after initiating CCl4 administration, body weights (g) in the low-fat diet (LFD), HFD + CCl4, and HFD + CCl4 + ELA (20 mg/kg) groups were 28.8 ± 1.2, 39.7 ± 2.9, and 38.8 ± 3.4, respectively. After 2 weeks of treatment, the values were 29.5 ± 1.7, 38.8 ± 3.3, and 36.9 ± 2.1, respectively, indicating a reduction in body weight in the HFD + CCl4 + ELA group relative to the HFD + CCl4 group. Serum alanine aminotransferase (ALT) was significantly lower in the HFD + CCl4 + ELA group (382.6 ± 60.5) than in the HFD + CCl4 group (599.8 ± 51.8) at this time point. Using an a priori normalization criterion—defined as the absence of a significant difference from the LFD control—these data indicate a marked attenuation of liver injury by ELA rather than complete normalization (Fig. 2).

Fig. 2.

Fig. 2

ELA-mediated effects on body weight changes and transaminase (AST and ALT) levels in HFD + CCl4-induced mice. Data are presented as mean ± SEM. ****p < 0.0001 or ***p < 0.001 vs. LFD group; *p < 0.05 vs. HFD + CCl4 group. CCl4 carbon tetrachloride, ELA elafibranor, AST aspartate aminotransferase, ALT alanine aminotransferase

On the last day of the experiment, after fasting, blood samples were collected from the posterior vena cava, and the left lateral lobe of the liver was harvested. Subsequently, the liver tissue was cut transversely from the middle section and immediately fixed in 10% neutral buffered formalin for histopathological examination. The fixed liver tissue was processed using a standard paraffin-embedding method and serially sectioned into 4-mm-thick sections. Liver fibrosis was observed using the Sirius Red staining technique [27]. Picrosirius Red staining was performed on deparaffinized tissue sections using a commercial kit (ab150681, Abcam, UK) according to the manufacturer’s instructions. Collagen fibers were visualized as red against a pale-yellow background.

Data preparation and implementation of algorithms

To obtain the WSIs from 19 glass slides, an Aperio ScanScope XT (Leica Biosystems, Nussloch, Germany) digital slide scanner was used to digitize the slides at 20 × magnification under brightfield illumination with a resolution of 0.5 μm/pixel. Subsequently, the acquired WSIs were divided into multiples of 120 × 120 and labeled as cropped patches with sizes of 32 × 32, 64 × 64, and 128 × 128 to apply various patch-based algorithms. The patches were classified into fibrotic tissue, normal tissue, and slide background images after a pathologist confirmed the presence or absence of fibrosis. Notably, the CNN model used in this study was implemented using the Keras framework [28]. Additionally, the hyperparameters, including learning rate, epoch, and batch size, required for model training, were optimized. All training processes were performed using an NVIDIA RTX 2080ti 11 GB Graphics Processing Unit (GPU). For each patch size (32 × 32, 64 × 64, and 128 × 128), training, validation, and testing steps were set up, and the dataset was randomly divided using an 8:1:1 ratio to learn the algorithms (Table 1).

Table 1.

Preparation of datasets for algorithm learning

Dataset Number of patch images
Size of 32 × 32 Size of 64 × 64 Size of 128 × 128
Fibrosis area Normal area Background Fibrosis area Normal area Background Fibrosis area Normal area Background
Training dataset 2952 2952 2952 1536 1536 1536 384 384 384
Validation dataset 369 369 369 192 192 192 48 48 48
Test dataset 369 369 369 192 192 192 48 48 48
Total 3690 3690 3690 1920 1920 1920 480 480 480

Number of patch images in training, validation, and test datasets for patch-based algorithm models

The structure of the algorithm models used in this study is shown in Fig. 3, and the implementation model summary is presented in Table 2. Data augmentation techniques, including random flip, rotation, and zoom, were used to expand the training data to increase data diversity. This approach generates an extensive amount of data, enabling algorithmic models to comprehensively learn different aspects of data distribution through various methods. The patch-based algorithms with sizes of 32 × 32, 64 × 64, and 128 × 128 formed three pairs of convolution and max-pooling layer blocks. For each pair, the number of filters in the convolution layer was gradually increased to 16, 32, and 64, respectively, to reduce the input image size. Additionally, the padding technique was used to match the size of the output feature map with that of the input image. Each convolution layer analyzed the characteristics of the image through the ReLU activation function. Furthermore, the fully connected layer was replaced with a dropout layer with a rate of 0.2 to reduce inter-layer connections. Subsequently, the layer was removed. The final output was normalized using the Softmax function. This enabled the classification of the patch images into the following three classes: normal tissue, fibrosis tissue, and slide background images. During training, the Adam (Adaptive Moment Estimation) optimizer algorithm was used to update the model parameters, with the learning rate set at 0.001. Moreover, categorical cross-entropy was used as the loss function for training, and each algorithm model was trained with a batch size of 32 for 40 epochs. The training was terminated when no further improvement was observed. The metric evaluated during training was accuracy [29].

Fig. 3.

Fig. 3

Convolution neural network architecture. WSI, whole slide images

Table 2.

Model summary in each patch-based algorithm model

Patch size of the 32 × 32 model
Layer type Output shape Parameters
Sequential (None, 32, 32, 3) 0
Rescaling (None, 32, 32, 3) 0
Conv2D (None, 32, 32, 16) 448
MaxPooling2D (None, 16, 16, 16) 0
Conv2D (None, 16, 16, 32) 4,640
MaxPooling2D (None, 8, 8, 32) 0
Conv2D (None, 8, 8, 64) 18,496
MaxPooling2D (None, 4, 4, 64) 0
Dropout (None, 4, 4, 64) 0
Flatten (None, 1,024) 0
Dense (None, 128) 131,200
Dense (None, 3) 387
Total parameters: 155,171
Trainable parameters: 155,171
Non-trainable parameters: 0
Patch size of the 64 × 64 model
Layer type Output shape Parameters
Sequential (None, 64, 64, 3) 0
Rescaling (None, 64, 64, 3) 0
Conv2D (None, 64, 64, 16) 448
MaxPooling2D (None, 32, 32, 16) 0
Conv2D (None, 32, 32, 32) 4,640
MaxPooling2D (None, 16, 16, 32) 0
Conv2D (None, 16, 16, 64) 18,496
MaxPooling2D (None, 8, 8, 64) 0
Dropout (None, 8, 8, 64) 0
Flatten (None, 4,096) 0
Dense (None, 128) 524,416
Dense (None, 3) 387
Total parameters: 548,387
Trainable parameters: 548,387
Non-trainable parameters: 0
Patch size of the 128 × 128 model
Layer type Output shape Parameters
Sequential (None, 128, 128, 3) 0
Rescaling (None, 128, 128, 3) 0
Conv2D (None, 128, 128, 16) 448
MaxPooling2D (None, 64, 64, 16) 0
Conv2D (None, 64, 64, 32) 4640
MaxPooling2D (None, 32, 32, 32) 0
Conv2D (None, 32, 32, 64) 18,496
MaxPooling2D (None, 16, 16, 64) 0
Dropout (None, 16, 16, 64) 0
Flatten (None, 16,384) 0
Dense (None, 128) 2,097,280
Dense (None, 3) 387
Total parameters: 2,121,251
Trainable parameters: 2,121,251
Non-trainable parameters: 0

After the training and validation steps, accuracy, precision, recall, and F1-score were utilized to evaluate the performance of the algorithm models using the test dataset [30]. Accuracy is defined as the ratio of correctly predicted instances to the total number of predictions (Eq. (1)).

Accuracy=TP+TN/(TP+TN+FP+FN) 1

Precision is the proportion of actual positive instances among the instances predicted as positive by the model. It evaluates the accuracy of the model’s positive predictions (Eq. (2)).

Precision=TP/(TP+FP) 2

Furthermore, recall is the proportion of actual positive instances that the model correctly predicted as positive. It evaluates how well the model identifies actual positives (Eq. (3)).

Recall=TP/(TP+FN) 3

The F1-score is calculated as the harmonic mean of precision and recall (Eq. (4)). In algorithm models, class imbalance problems can occur, making it difficult to adequately evaluate the model’s performance based on accuracy alone. Moreover, precision and recall are in a trade-off relationship. Therefore, the F1-score, which is used to assess the balance between these two metrics, can provide a more comprehensive performance evaluation than simply evaluating the model based on accuracy.

F1Score=2×Precision×RecallPrecision+Recall 4

WSI analysis of the predicted values of algorithms and the ground truth annotation by a pathologist

The developed algorithms were compared with the fibrotic regions directly annotated by a pathologist through WSI analysis to assess their similarity. For each algorithm, the predicted values were generated by inputting patch images. The predicted values are presented in Table 3. Furthermore, the predicted values were generated in patch form; therefore, they were stitched into WSI form based on their labeled coordinates. For the stitched WSIs, the fibrotic areas were calculated based on the pixel values of each algorithm (0.499 μm/pixel).

Table 3.

Predicted values of patch-based algorithm models

Patch-based models Number of patch images WSI
Fibrosis area Normal area
32 × 32 model 118,612 1,994,808 19
64 × 64 model 56,300 489,400 19
128 × 128 model 22,505 118,306 19

WSI whole slide image

Number of output patch images of fibrotic tissue, normal tissue, and background images predicted by each algorithm

The fibrotic area for a single WSI was calculated as follows (Eq. (5)):

Absolute fibrotic area(μm2)=(patch size×0.499μm)2×Number of fibrosis patches 5

The fibrotic ratio (%) relative to each actual liver tissue was also calculated. Actual liver tissue refers to the sum of fibrotic tissue and normal tissue. The fibrotic ratio for a single WSI was calculated as follows (Eq. (6)):

Relative fibrotic ratio(\%)=Number of fibrosis patchesNumber of actual liver tissue patches×100 6

Furthermore, to compare with the values predicted by each algorithm, a pathologist directly annotated the liver fibrosis area on 19 WSIs for each mouse using the Aperio ImageScope area analysis program. This enabled the calculation of the fibrotic area (μm2) in the actual liver tissue. We considered the fibrotic area annotated by the pathologist on the WSIs as the ground truth. Subsequently, Spearman’s correlation coefficient, a non-parametric correlation analysis, was used to evaluate the similarity between the algorithm-predicted values and the pathologist’s ground truth.

Statistical analyses by group

An additional comparative analysis was performed among treatment groups. An ordinary one-way analysis of variance (ANOVA) followed by multiple comparison tests was conducted to evaluate differences among the vehicle control, HFD + CCl4, and HFD + CCl4 + ELA groups. Results are presented as means ± standard deviations. Statistical significance was defined as p < 0.0001. All analyses were performed using Prism 10 (GraphPad Software, Inc., CA, USA).

Results

Results of algorithm model tests

The algorithms for the patches with 32 × 32, 64 × 64, and 128 × 128 sizes were developed through training, validation, and testing steps using an 8:1:1 ratio. During this process, the batch size of each algorithm was set at 32. Training and validation were conducted for 40 epochs. The 32 × 32 patch size algorithm achieved a validation accuracy of 98.19% (Fig. 4a), while the 64 × 64 and 128 × 128 algorithms reached validation accuracies of 99.13% (Fig. 4b) and 98.61% (Fig. 4c), respectively. Additionally, all three algorithms showed validation losses below 0.1.

Fig. 4.

Fig. 4

Results of each model’s accuracy. Algorithm accuracy and loss graphs for training (blue) and validation (orange) based on 40 epochs and 32 batch sizes

Furthermore, confusion matrices were used to analyze the relationship between actual and predicted values to evaluate the performance of the algorithm classification models. For each algorithm, model prediction tests were performed using test datasets with patch sizes of 32 × 32, 64 × 64, and 128 × 128. The results revealed that the 32 × 32 algorithm had one error among the 369 fibrotic patches (Fig. 5a), the 64 × 64 algorithm recorded three errors among the 192 fibrotic patches (Fig. 5b), and the 128 × 128 algorithm showed two errors among the 48 fibrotic patches (Fig. 5c).

Fig. 5.

Fig. 5

Results of the confusion matrix. Confusion matrix using the test dataset for each algorithm

Representative examples of prediction errors are shown in Fig. 6. A background patch was misclassified as a fibrotic tissue patch, representing a false positive (Fig. 6d). A fibrotic tissue patch was misclassified as a normal tissue patch, representing a false negative (Fig. 6e). Background region patches were also misclassified as normal tissue patches (Fig. 6f).

Fig. 6.

Fig. 6

Error cases during the test steps. a–c Representative images from the vehicle control, CCl4, and CCl4 + ELA groups, respectively, d background area misclassified as fibrotic tissue, e fibrotic tissue misclassified as normal tissue, f background areas misclassified as normal tissues

All patch-based algorithms achieved precision and recall above 0.98 and 0.96, respectively, for fibrotic findings. Furthermore, the F1-score ranges from 0 to 1, with values closer to 1 indicating better model performance. Our trained patch-based algorithms, comprising three classes in total, achieved F1 scores above 0.96 (Table 4).

Table 4.

Deep learning results in each patch-based algorithm model

32 × 32 model Classification report Accuracy
Precision Recall F1-score
Fibrosis area 1.00 1.00 1.00 0.98
Normal area 0.95 1.00 0.97
Background 1.00 0.95 0.97
64 × 64 model Classification report Accuracy
Precision Recall F1-score
Fibrosis area 0.99 0.98 0.99 0.99
Normal area 0.97 0.99 0.98
Background 1.00 0.99 0.99
128 × 128 model Classification report Accuracy
Precision Recall F1-score
Fibrosis area 0.98 0.96 0.97 0.97
Normal area 0.94 1.00 0.97
Background 0.98 0.94 0.96

Values in developed algorithms and ground truth annotation

A visual comparison of the actual values (annotations) of fibrotic regions directly annotated by a pathologist on WSIs with the predicted values of the three patch-based algorithms is shown in Fig. 7. In this figure, the pathologist’s actual values and predicted values of the 64 × 64 algorithm appear visually most similar; however, this was because of the thickness of the digital pen tool used during the annotation process. Therefore, the absolute area (μm2) of the fibrotic regions was calculated for a more accurate comparison (Fig. 8). Notably, the predicted fibrotic absolute area values of the 32 × 32 algorithm were most similar to the pathologist’s actual values. Moreover, the similarity of the predicted values generated by the three algorithms was evaluated based on the actual values of the fibrotic regions annotated by the pathologist. The Spearman’s correlation coefficient, a non-parametric correlation, was used. The results revealed that the 32 × 32 algorithm had the highest correlation (Spearman’s r = 0.9609) with the pathologist’s actual values, whereas the 64 × 64 and 128 × 128 algorithms had correlation coefficients of 0.9298 and 0.8930, respectively (Fig. 9). The 95% confidence intervals for each algorithm were as follows: 0.8964–0.9856, 0.8186–0.9738, and 0.7315–0.9596 for the 32 × 32, 64 × 64, and 128 × 128 algorithms, respectively. Additionally, the p-values for all algorithms were less than 0.0001 (p < 0.05).

Fig. 7.

Fig. 7

Fibrosis area in whole slide image analysis. The fibrosis area is directly annotated by a pathologist on the whole slide images, and the stitch represents the predicted values of the fibrosis area by each algorithm

Fig. 8.

Fig. 8

Comparison of the absolute fibrosis area (μm2) on the whole slide image. CON vehicle control, CCl4 carbon tetrachloride, ELA elafibranor

Fig. 9.

Fig. 9

Spearman’s correlation of fibrosis area between the predicted values of each algorithm and annotation

Group mean of fibrosis ratio compared with controls

In this study, comparisons between groups were also performed. The ratio (%) of the fibrotic area to the actual liver tissue area was calculated using five vehicle control mice, nine positive control mice (HFD + CCl4 group), and five test mice (HFD + CCl4 + ELA group). The results of the ordinary one-way ANOVA followed by multiple comparison tests revealed that the fibrotic area ratio between the vehicle control and HFD + CCl4 groups significantly differed between the pathologist’s actual values and predicted values from all developed patch-based algorithms (p < 0.0001). Furthermore, the fibrotic area ratio between the HFD + CCl4 and HFD + CCl4 + ELA groups significantly differed between the pathologist’s actual values and predicted values from all developed patch-based algorithms (p < 0.0001 or p = 0.0002) (Fig. 10).

Fig. 10.

Fig. 10

Differences in mean value between groups. CON vehicle control, CCl4 carbon tetrachloride, ELA elafibranor. Statistical significance was calculated using ordinary one-way ANOVA. **** indicates p < 0.0001; *** denotes p = 0.0002

Discussion

Liver fibrosis is a key histopathological marker of chronic hepatic injury caused by toxic agents, including CCl4 [31]. In preclinical toxicology studies, accurate and reproducible quantification of fibrosis is critical for assessing the severity of toxic damage and the therapeutic efficacy of candidate compounds [32]. Incorporating advanced image analysis methods, particularly those powered by AI, can improve the consistency and objectivity of fibrosis assessment, thereby enhancing toxicological evaluations and facilitating drug development.

In this study, three patch‑based classification algorithms (32 × 32, 64 × 64, and 128 × 128) were developed to assess liver fibrosis patterns at multiple scales. All models demonstrated high accuracy and strong predictive performance; however, the 32 × 32 algorithm generated fibrotic area estimates that most closely matched the pathologist’s annotations. Liver fibrosis was confirmed using the Sirius Red staining technique, and the developed algorithm models were evaluated using WSI image analysis based on the collagen content of liver tissue. Using the proposed method, a board-certified veterinary pathologist evaluated hepatic fibrosis in nine mice from the CCl4-induced positive control group, and the severity was assessed as ranging from slight to moderate. Fibrosis ratios derived from pathologist annotations ranged from 4.1% to 6.7%, while those predicted by the 32 × 32 algorithm ranged from 6.3% to 10.2%. In animals classified as having moderate fibrosis, annotation-based fibrosis ratios exceeded approximately 5.5%, and algorithmic estimates were above 7.8%, showing comparable increases in fibrosis severity between the two assessment approaches (Table 5). These findings indicate that the proposed algorithm can sensitively capture gradations of fibrosis severity within the slight-to-moderate range. Moreover, it may help reduce inter-observer variability and fatigue among pathologists during large-scale histopathological assessments, thereby enabling more consistent quantification of fibrosis [22, 33].

Table 5.

Comparison of pathologist annotation and algorithm-based fibrosis ratios in the HFD + CCl4 group

Animals Fibrosis ratio (%) Severity
Annotation 32 × 32 model
CCl4 1 4.9 7.2 Slight
CCl4 2 5.4 7.8 Moderate
CCl4 3 5.0 7.4 Slight
CCl4 4 5.5 6.9 Slight
CCl4 5 4.9 6.3 Slight
CCl4 6 6.7 10.2 Moderate
CCl4 7 5.8 9.0 Moderate
CCl4 8 4.1 7.1 Slight
CCl4 9 5.4 8.8 Moderate

Although hepatic fibrosis can be readily visualized using conventional staining methods, including Sirius Red or α-SMA [34, 35], quantitative assessment across entire whole-slide images at high magnification remains limited by the lack of automated analysis tools [36, 37]. The proposed AI-based approach enables objective, large-scale quantification of fibrosis by learning from expert pathologist annotations, thereby extending expert-level evaluation to a reproducible and standardized framework. Therefore, this underscores its potential utility in preclinical toxicology, where quantitative, reproducible, and efficient evaluation of histopathological lesions is critical for drug safety assessment and mechanistic research [38].

The test group that received additional ELA showed a decreased liver fibrosis ratio (%) compared with the positive control group, where liver fibrosis was induced using CCl4 (Fig. 10). The positive control group had a high distribution of fibrosis, leading to an enlarged liver. Furthermore, in the test group treated with ELA, necrotic features appeared due to the antifibrotic effect, and fibrosis showed an intermittently disrupted form. Consequently, the relative fibrosis ratio (%) was most similar between the pathologist’s actual values and the predicted values of the 32 × 32 algorithm (Fig. 11). These findings are consistent with those of previous reports demonstrating that ELA improves hepatic steatosis and inflammation and delays the progression of liver fibrosis in preclinical and phase 2 clinical trials [39]. Although ELA has been associated with mild and reversible elevations in liver enzymes at higher doses, it has also demonstrated hepatoprotective and antifibrotic effects in models of metabolic dysfunction–associated steatohepatitis [40–42]. In this study, ELA was administered at 20 mg/kg [25, 26] starting 2 weeks after CCl4 exposure. Under this regimen, the HFD + CCl4 + ELA group exhibited decreased serum ALT levels compared with the HFD + CCl4 group, indicating no apparent hepatotoxicity under our experimental conditions.

Fig. 11.

Fig. 11

Comparison of the relative fibrosis area (%) on the whole slide image. CON vehicle control, CCl4 carbon tetrachloride, ELA elafibranor

Each lesion has its unique morphological pattern in histopathological examination. Particularly, fibrotic tissue is characterized by irregular and atypical patterns rather than linear or spherical shapes [43], making it challenging to train the algorithms to learn all these components [44]. Each algorithm tended to primarily recognize the red portions of patch images from slide background areas (artifacts such as the red dye of this staining method (Fig. 6d) or blood or keratin (Fig. 6f)) as fibrosis or normal tissue. A pathologist directly judged and considered artifacts when annotating; however, the trained algorithms lacked clear criteria for distinguishing between normal (or slide background images) and fibrotic regions based on the degree of redness (Fig. 6e). Therefore, future studies should include additional training with various normal tissue WSIs to enable AI algorithms to learn artifacts.

In the case of mouse models, selecting an appropriate patch size was important to identify the optimal combination with the CNN architecture because of the small size of liver tissue. Relatively small patch sizes may miss important lesion information, whereas relatively large patch sizes may increase computational complexity and require more GPU memory. Consequently, the predicted values of the algorithm using the smallest patch size (32 × 32) included significantly fewer non-fibrotic areas than those using relatively larger patch sizes (64 × 64 or 128 × 128). Notably, the 128 × 128 algorithm included the largest proportion of non-fibrotic areas when predicting fibrotic regions, resulting in differences between the vehicle control and test groups. Therefore, this study’s results indicate that using a smaller patch size (32 × 32) enables more precise capture of fibrotic regions in mouse liver tissue through digital image analysis, thereby reducing the inclusion of non‑fibrotic areas in the prediction (Fig. 12).

Fig. 12.

Fig. 12

Comparison of fibrosis between the algorithms. A combination of digital images of fibrosis labeled with Sirius Red-staining and algorithmic predictions for each cropped patch. Subsequently, a complementary color effect is used to enhance visual distinction for easy identification by the naked eye. Figure was created using Procreate (version 5.3.15 build, Savage Interactive, iPadOS)

AI algorithms are generally categorized into classification and segmentation models. Segmentation models extract regions of interest at the pixel level but are computationally intensive and require large datasets for large-scale images. In contrast, classification models assign labels at the image level, enabling simpler and faster computation. Our patch-based approach combines the strengths of both methods [45]. WSI images of liver tissue were divided into small patches, which the algorithms classified as fibrotic, normal tissue, or slide background. Subsequently, these classified patches were reassembled into the original WSI, enabling more detailed characterization of fibrosis. This process allowed precise calculation of the absolute fibrotic area (μm2) and its ratio (%) to total liver tissue. Colorimetric analysis using tools, including ImageJ, can measure collagen staining intensity; however, threshold variations from slide to slide and artifacts in non-tissue regions frequently reduce consistency. Moreover, such methods are limited to selected areas rather than whole-slide evaluation [46, 47]. In contrast, the proposed AI-based approach learns diverse staining patterns and artifacts, enabling more robust and consistent quantification across entire slides.

Other preclinical studies have quantified fibrotic tissue stained with this method using deep learning. One study semi-quantitatively analyzed liver fibrosis in female mice using images with a size of 1024 × 1024. Another study compared digital image analysis of patch images at × 10 and × 40 magnifications with fibrotic regions directly annotated by pathologists using machine learning algorithms. The results showed that fibrotic patch images at × 40 magnification strongly correlated (Spearman r = 0.939) with pathologists’ annotations, demonstrating the potential of deep learning as a decision-support tool [48].

Furthermore, other staining methods, including α-SMA immunohistochemistry and Masson’s Trichrome, have been applied for the quantitative assessment of fibrosis using AI-based image analysis [49, 50]. With the establishment of appropriate reference standards and annotation protocols, incorporating these staining techniques is expected to further enhance the robustness and applicability of automated fibrosis evaluation [51–53].

Compared with previous studies that applied deep learning–based semi-quantitative analyses to relatively large image sizes, this study developed high-precision classification models using multiple patch sizes (32 × 32, 64 × 64, and 128 × 128). This multi-scale approach not only achieved higher accuracy but also demonstrated substantial practical utility by automatically quantifying the antifibrotic effects of the investigational drug ELA in a preclinical toxicology model. These findings suggest that the proposed models can serve as reliable and efficient tools for liver fibrosis assessment in toxicological research and drug development, providing an objective and reproducible platform for pathological evaluation.

Although this study was conducted under a single experimental design using CCl4‑induced liver fibrosis in C57BL/6 mice, it serves as an initial proof of concept for the proposed patch‑based algorithms. However, potential bias inherent to the patch-based data classification approach remains a limitation of this study, as even subtle variations in staining intensity or tissue processing within the same slide could influence model performance [54, 55]. To minimize such bias, future research should implement data partitioning at the slide or individual animal level. To broaden their applicability, future studies should include external validation using additional animal models, incorporation of multi‑institutional datasets, and evaluation under varied slide‑scanning and tissue‑preparation protocols. Leveraging transfer learning, these models can be readily adapted to diverse experimental settings while maintaining robust performance even with relatively small datasets, underscoring their strong potential for broad application in preclinical toxicology.

The developed algorithm models could learn distinct fibrotic regions using this method compared with other staining methods. Additionally, they enabled sophisticated quantitative analysis using patch images with a size of at least 32 × 32. The fibrotic patch-based model in this study can be implemented with basic parameters and a short time while achieving high accuracy, unlike segmentation models that require complex parameters. These are encouraging results for the automatic quantitative evaluation of histopathological lesions. Therefore, pathologists can be provided with consistent quantitative thresholds between the normal range and fibrosis using these AI algorithms.

While this study focused on developing AI models for fibrosis analysis, future research will aim to expand this approach to other hepatic pathologies, including fatty liver and steatohepatitis. By acquiring comprehensive datasets that capture lipid droplet morphology and associated inflammatory features, it will be possible to establish integrated AI models capable of simultaneously analyzing multiple histopathological findings of liver injury [56, 57]. Such models would further enhance the clinical and toxicological relevance of AI-assisted pathology.

In conclusion, our results demonstrate the potential of integrating AI tools into pathology workflows to standardize and enhance liver fibrosis assessment, with possible extension to clinical pathology practice.

Supplementary Information

Below is the link to the electronic supplementary material.

Acknowledgements

Assistance from Editage (www.editage.co.kr) regarding language editing is acknowledged.

Author contributions

Jae‑Woo Cho oversaw project administration and funding acquisition, contributed to conceptualization and investigation, performed formal analysis, and provided supervision. Jin‑Hee Lee conceived the study, conducted the investigation, curated the data, performed formal analysis and visualization, and drafted the original manuscript. Myung‑Hwa Yang and Gyeongjin Han performed formal analysis, developed the methodology, and curated the data. Won Hoon Jung and Myung Ae Bae provided resources and supervision. Ji‑Seok Han and Tae‑Sung Koo curated the data and provided supervision. All authors read and commented on the manuscript.

Funding

This work was supported by the Korea Institute of Toxicology, KIT, Republic of Korea (grant number 2710086919).

Data availability

The datasets used and/or analyzed during the current study are available from the corresponding author upon reasonable request.

Declarations

Conflict of interest

The authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper.

Ethical approval

All animal experimental procedures were performed in accordance with relevant guidelines and regulations. Approval was granted by the Institutional Animal Care and Use Committee of the Korea Research Institute of Chemical Technology (Daejeon, Republic of Korea) (Permit Number: 2020-7A-02–02).

Consent to participate

Not applicable. This study did not involve human participants.

Consent to publish

Not applicable. This manuscript does not contain data from any individual person.

Footnotes

Publisher's Note

Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

References

  • 1.Rinella ME, Lazarus JV, Ratziu V, Francque SM, Sanyal AJ, Kanwal F, Romero D, Abdelmalek MF, Anstee QM, Arab JP, Arrese M, Bataller R, Beuers U, Boursier J, Bugianesi E, Byrne CD, Castro Narro GE, Chowdhury A, Cortez-Pinto H, Cryer DR, Cusi K, El-Kassas M, Klein S, Eskridge W, Fan J, Gawrieh S, Guy CD, Harrison SA, Kim SU, Koot BG, Korenjak M, Kowdley KV, Lacaille F, Loomba R, Mitchell-Thain R, Morgan TR, Powell EE, Roden M, Romero-Gómez M, Silva M, Singh SP, Sookoian SC, Spearman CW, Tiniakos D, Valenti L, Vos MB, Wong VW, Xanthakos S, Yilmaz Y, Younossi Z, Hobbs A, Villota-Rivas M, Newsome PN, NAFLD Nomenclature Consensus Group (2023) A multisociety Delphi consensus statement on new fatty liver disease nomenclature. Hepatology 78:1966–1986. 10.1097/HEP.0000000000000520 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2.Heyens LJM, Busschots D, Koek GH, Robaeys G, Francque S (2021) Liver fibrosis in non-alcoholic fatty liver disease: from liver biopsy to non-invasive biomarkers in diagnosis and treatment. Front Med (Lausanne) 8:615978. 10.3389/fmed.2021.615978 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3.Younossi ZM, Henry L (2024) Understanding the burden of nonalcoholic fatty liver disease: time for action. Diabetes Spectr 37:9–19. 10.2337/dsi23-0010 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.Roeb E (2021) Non-alcoholic fatty liver diseases: current challenges and future directions. Ann Transl Med 9:726. 10.21037/atm-20-3760 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Sumida Y, Yoneda M (2018) Current and future pharmacological therapies for NAFLD/NASH. J Gastroenterol 53:362–376. 10.1007/s00535-017-1415-1 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Parasuraman S (2011) Toxicological screening J Pharmacol Pharmacother 2:74–79. 10.4103/0976-500X.81895 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.Crissman JW, Goodman DG, Hildebrandt PK, Maronpot RR, Prater DA, Riley JH, Seaman WJ, Thake DC (2004) Best practices guideline: toxicologic histopathology. Toxicol Pathol 32:126–131. 10.1080/01926230490268756 [DOI] [PubMed] [Google Scholar]
  • 8.Homeyer A, Geissler C, Schwen LO, Zakrzewski F, Evans T, Strohmenger K, Westphal M, Bulow RD, Kargl M, Karjauv A, Munne-Bertran I, Retzlaff CO, Romero-Lopez A, Soltysinski T, Plass M, Carvalho R, Steinbach P, Lan YC, Bouteldja N, Haber D, Rojas-Carulla M, Vafaei Sadr A, Kraft M, Kruger D, Fick R, Lang T, Boor P, Muller H, Hufnagl P, Zerbe N (2022) Recommendations on compiling test datasets for evaluating artificial intelligence solutions in pathology. Mod Pathol 35:1759–1769. 10.1038/s41379-022-01147-y [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.Cooper LA, Carter AB, Farris AB, Wang F, Kong J, Gutman DA, Widener P, Pan TC, Cholleti SR, Sharma A, Kurc TM, Brat DJ, Saltz JH (2012) Digital pathology: data-intensive frontier in medical imaging: Health-information sharing, specifically of digital pathology, is the subject of this paper which discusses how sharing the rich images in pathology can stretch the capabilities of all otherwise well-practiced disciplines. Proc IEEE Inst Electr Electron Eng 100:991–1003. 10.1109/JPROC.2011.2182074 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.Melo RCN, Raas MWD, Palazzi C, Neves VH, Malta KK, Silva TP (2019) Whole slide imaging and its applications to histopathological studies of liver disorders. Front Med (Lausanne) 6:310. 10.3389/fmed.2019.00310 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11.Kumar N, Gupta R, Gupta S (2020) Whole slide imaging (WSI) in pathology: current perspectives and future directions. J Digit Imaging 33:1034–1040. 10.1007/s10278-020-00351-z [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Kiran N, Sapna F, Kiran F, Kumar D, Raja F, Shiwlani S, Paladini A, Sonam F, Bendari A, Perkash RS, Anjali F, Varrassi G (2023) Digital pathology: transforming diagnosis in the digital age. Cureus 15:e44620. 10.7759/cureus.44620 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.Cai L, Gao J, Zhao D (2020) A review of the application of deep learning in medical image classification and segmentation. Ann Transl Med 8:713. 10.21037/atm.2020.02.44 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.Sarvamangala DR, Kulkarni RV (2022) Convolutional neural networks in medical image understanding: a survey. Evol Intell 15:1–22. 10.1007/s12065-020-00540-3 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Liu Y, Kohlberger T, Norouzi M, Dahl GE, Smith JL, Mohtashamian A, Olson N, Peng LH, Hipp JD, Stumpe MC (2019) Artificial intelligence-based breast cancer nodal metastasis detection: insights into the black box for pathologists. Arch Pathol Lab Med 143:859–868. 10.5858/arpa.2018-0147-OA [DOI] [PubMed] [Google Scholar]
  • 16.McGenity C, Clarke EL, Jennings C, Matthews G, Cartlidge C, Freduah-Agyemang H, Stocken DD, Treanor D (2024) Artificial intelligence in digital pathology: a systematic review and meta-analysis of diagnostic test accuracy. NPJ Digit Med 7:114. 10.1038/s41746-024-01106-8 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Rudmann D, Albretsen J, Doolan C, Gregson M, Dray B, Sargeant A, O’Shea DD, Kuklyte J, Power A, Fitzgerald J (2021) Using deep learning artificial intelligence algorithms to verify N-nitroso-N-methylurea and urethane positive control proliferative changes in Tg-RasH2 mouse carcinogenicity studies. Toxicol Pathol 49:938–949. 10.1177/0192623320973986 [DOI] [PubMed] [Google Scholar]
  • 18.Popa SL, Ismaiel A, Cristina P, Cristina M, Chiarioni G, David D, Dumitrascu DL (2021) Non-alcoholic fatty liver disease: implementing complete automated diagnosis and staging. A systematic review. Diagnostics (Basel) 11:1078. 10.3390/diagnostics11061078 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.Liang L, Liu M, Sun W (2017) A deep learning approach to estimate chemically-treated collagenous tissue nonlinear anisotropic stress-strain responses from microscopy images. Acta Biomater 63:227–235. 10.1016/j.actbio.2017.09.025 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.Heinemann F, Birk G, Schoenberger T, Stierstorfer B (2018) Deep neural network based histological scoring of lung fibrosis and inflammation in the mouse model system. PLoS ONE 13:e0202708. 10.1371/journal.pone.0202708 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.Heinemann F, Birk G, Stierstorfer B (2019) Deep learning enables pathologist-like scoring of NASH models. Sci Rep 9:18454. 10.1038/s41598-019-54904-6 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Grignaffini F, Barbuto F, Troiano M, Piazzo L, Simeoni P, Mangini F, De Stefanis C, Onetti Muda A, Frezza F, Alisi A (2024) The use of artificial intelligence in the liver histopathology field: a systematic review. Diagnostics (Basel) 14:388. 10.3390/diagnostics14040388 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23.Damiris K, Tafesh ZH, Pyrsopoulos N (2020) Efficacy and safety of anti-hepatic fibrosis drugs. World J Gastroenterol 26:6304–6321. 10.3748/wjg.v26.i41.6304 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24.Dong S, Chen QL, Song YN, Sun Y, Wei B, Li XY, Hu YY, Liu P, Su SB (2016) Mechanisms of CCl4-induced liver fibrosis with combined transcriptomic and proteomic analysis. J Toxicol Sci 41:561–572. 10.2131/jts.41.561 [DOI] [PubMed] [Google Scholar]
  • 25.Briand F, Maupoint J, Brousseau E, Breyner N, Bouchet M, Costard C, Leste-Lasserre T, Petitjean M, Chen L, Chabrat A, Richard V, Burcelin R, Dubroca C, Sulpice T (2021) Elafibranor improves diet-induced nonalcoholic steatohepatitis associated with heart failure with preserved ejection fraction in Golden Syrian hamsters. Metabolism 117:154707. 10.1016/j.metabol.2021.154707 [DOI] [PubMed] [Google Scholar]
  • 26.Briand F, Heymes C, Bonada L, Angles T, Charpentier J, Branchereau M, Brousseau E, Quinsat M, Fazilleau N, Burcelin R, Sulpice T (2020) A 3-week nonalcoholic steatohepatitis mouse model shows elafibranor benefits on hepatic inflammation and cell death. Clin Transl Sci 13:529–538. 10.1111/cts.12735 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27.Yu Y, Wang J, Ng CW, Ma Y, Mo S, Fong ELS, Xing J, Song Z, Xie Y, Si K, Wee A, Welsch RE, So PT, Yu H (2018) Deep learning enables automated scoring of liver fibrosis stages. Sci Rep 8:16016. 10.1038/s41598-018-34300-2 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28.Kim HJ, Baek EB, Hwang JH, Lim M, Jung WH, Bae MA, Son HY, Cho JW (2023) Application of convolutional neural network for analyzing hepatic fibrosis in mice. J Toxicol Pathol 36:21–30. 10.1293/tox.2022-0066 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29.Alzubaidi L, Zhang J, Humaidi AJ, Al-Dujaili A, Duan Y, Al-Shamma O, Santamaria J, Fadhel MA, Al-Amidie M, Farhan L (2021) Review of deep learning: concepts, CNN architectures, challenges, applications, future directions. J Big Data 8:53. 10.1186/s40537-021-00444-8 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30.Hicks SA, Strumke I, Thambawita V, Hammou M, Riegler MA, Halvorsen P, Parasa S (2022) On evaluation metrics for medical applications of artificial intelligence. Sci Rep 12:5979. 10.1038/s41598-022-09954-8 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31.Yu C, Wang F, Jin C, Wu X, Chan WK, McKeehan WL (2002) Increased carbon tetrachloride-induced liver injury and fibrosis in FGFR4-deficient mice. Am J Pathol 161:2003–2010. 10.1016/S0002-9440(10)64478-1 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32.Zheng WV, Li Y, Cheng X, Xu Y, Zhou T, Li D, Xiong Y, Wang S, Chen Z (2022) Uridine alleviates carbon tetrachloride-induced liver fibrosis by regulating the activity of liver-related cells. J Cell Mol Med 26:840–854. 10.1111/jcmm.17131 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33.Shabanian M, Taylor Z, Woods C, Bernieh A, Dillman J, He L, Ranganathan S, Picarsic J, Somasundaram E (2025) Liver fibrosis classification on trichrome histology slides using weakly supervised learning in children and young adults. J Pathol Inform 16:100416. 10.1016/j.jpi.2024.100416 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34.Lo RC, Kim H (2017) Histopathological evaluation of liver fibrosis and cirrhosis regression. Clin Mol Hepatol 23:302–307. 10.3350/cmh.2017.0078 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 35.Akpolat N, Yahsi S, Godekmerdan A, Yalniz M, Demirbag K (2005) The value of alpha-SMA in the evaluation of hepatic fibrosis severity in hepatitis B infection and cirrhosis development: a histopathological and immunohistochemical study. Histopathology 47:276–280. 10.1111/j.1365-2559.2005.02226.x [DOI] [PubMed] [Google Scholar]
  • 36.Rizzo PC, Girolami I, Marletta S, Pantanowitz L, Antonini P, Brunelli M, Santonicco N, Vacca P, Tumino N, Moretta L, Parwani A, Satturwar S, Eccher A, Munari E (2022) Technical and diagnostic issues in whole slide imaging published validation studies. Front Oncol 12:918580. 10.3389/fonc.2022.918580 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 37.Serdjebi C, Bertotti K, Huang P, Wei G, Skelton-Badlani D, Leclercq IA, Barbes D, Lepoivre B, Popov YV, Julé Y (2022) Automated whole slide image analysis for a translational quantification of liver fibrosis. Sci Rep 12:17935. 10.1038/s41598-022-22902-w [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 38.Jaume G, de Brot S, Song AH, Williamson DFK, Oldenburg L, Zhang A, Chen RJ, Asin J, Blatter S, Dettwiler M, Goepfert C, Grau-Roma L, Soto S, Keller SM, Rottenberg S, Del-Pozo J, Pettit R, Le LP, Mahmood F (2024) Deep learning-based modeling for preclinical drug safety assessment. bioRxiv. 10.1101/2024.07.20.604430 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 39.van den Hoek AM, Verschuren L, Caspers MPM, Worms N, Menke AL, Princen HMG (2021) Beneficial effects of elafibranor on NASH in E3L.CETP mice and differences between mice and men. Sci Rep 11:5050. 10.1038/s41598-021-83974-8 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 40.National Institute of Diabetes and Digestive and Kidney Diseases (2015) Elafibranor. LiverTox: clinical and research information on drug-induced liver injury. National Institute of Health, Bethesda. https://pubmed.ncbi.nlm.nih.gov/31643176/ [Google Scholar]
  • 41.Schattenberg JM, Pares A, Kowdley KV, Heneghan MA, Caldwell S, Pratt D, Bonder A, Hirschfield GM, Levy C, Vierling J, Jones D, Tailleux A, Staels B, Megnien S, Hanf R, Magrez D, Birman P, Luketic V (2021) A randomized placebo-controlled trial of elafibranor in patients with primary biliary cholangitis and incomplete response to UDCA. J Hepatol 74:1344–1354. 10.1016/j.jhep.2021.01.013 [DOI] [PubMed] [Google Scholar]
  • 42.Roth JD, Veidal SS, Fensholdt LKD, Rigbolt KTG, Papazyan R, Nielsen JC, Feigh M, Vrang N, Young M, Jelsing J, Adorini L, Hansen HH (2019) Combined obeticholic acid and elafibranor treatment promotes additive liver histological improvements in a diet-induced ob/ob mouse model of biopsy-confirmed NASH. Sci Rep 9:9046. 10.1038/s41598-019-45178-z [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 43.Chobert MN, Couchie D, Fourcot A, Zafrani ES, Laperche Y, Mavier P, Brouillet A (2012) Liver precursor cells increase hepatic fibrosis induced by chronic carbon tetrachloride intoxication in rats. Lab Invest 92:135–150. 10.1038/labinvest.2011.143 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 44.Masugi Y, Abe T, Tsujikawa H, Effendi K, Hashiguchi A, Abe M, Imai Y, Hino S, Hige K, Kawanaka M, Yamada G, Kage M, Korenaga M, Hiasa Y, Mizokami M, Sakamoto M (2018) Quantitative assessment of liver fibrosis reveals a nonlinear association with fibrosis stage in nonalcoholic fatty liver disease. Hepatol Commun 2:58–68. 10.1002/hep4.1121 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 45.Greeley C, Holder L, Nilsson EE, Skinner MK (2024) Scalable deep learning artificial intelligence histopathology slide analysis and validation. Sci Rep 14:26748. 10.1038/s41598-024-76807-x [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 46.Courtoy GE, Leclercq I, Froidure A, Schiano G, Morelle J, Devuyst O, Huaux F, Bouzin C (2020) Digital image analysis of Picrosirius Red staining: a robust method for multi-organ fibrosis quantification and characterization. Biomolecules 10:1585. 10.3390/biom10111585 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 47.Facchin C, Certain A, Yoganathan T, Delacroix C, Arevalo Garcia A, Gaillard F, Lenoir O, Tharaux PL, Tavitian B, Balvay D (2022) Fiber-ML, an open-source supervised machine learning tool for quantification of fibrosis in tissue sections. Am J Pathol 192:783–793. 10.1016/j.ajpath.2022.01.013 [DOI] [PubMed] [Google Scholar]
  • 48.Ramot Y, Deshpande A, Morello V, Michieli P, Shlomov T, Nyska A (2021) Microscope-based automated quantification of liver fibrosis in mice using a deep learning algorithm. Toxicol Pathol 49:1126–1133. 10.1177/01926233211003866 [DOI] [PubMed] [Google Scholar]
  • 49.Hillsley A, Santos JE, Rosales AM (2021) A deep learning approach to identify and segment alpha-smooth muscle actin stress fiber positive cells. Sci Rep 11:21855. 10.1038/s41598-021-01304-4 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 50.Sánchez-Jaramillo EA, Maldonado-Suárez LM, Torres-Rocha JF, Gutiérrez-Pérez C, Rojas-Valencia ML, Ramírez-Clavijo SR, Sánchez-Muñoz JF, Uribe-García M, Aristizábal-Gutiérrez M, Ochoa-Rojas A (2022) Automated computer-assisted image analysis for the fast quantification of renal fibrosis in Masson’s trichrome-stained sections. Diagnostics (Basel) 12:1906. 10.3390/diagnostics12081906 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 51.Song L, Zhang D, Wang H, Xia X, Huang W, Gonzales J, Via LE, Wang D (2024) Automated quantitative assay of fibrosis characteristics in tuberculosis granulomas. Front Microbiol 14:1301141. 10.3389/fmicb.2023.1301141 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 52.Popa SL, Ismaiel A, Abenavoli L, Padureanu AM, Dita MO, Bolchis R, Munteanu MA, Brata VD, Pop C, Bosneag A, Dumitrascu DI, Barsan M, David L (2023) Diagnosis of liver fibrosis using artificial intelligence: a systematic review. Medicina (Kaunas) 59:992. 10.3390/medicina59050992 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 53.Yegin EG, Yegin K, Karatay E, Kombak EF, Tuney D, Ataizi-Celikel C, Ozdogan OC (2015) Quantitative assessment of liver fibrosis by digital image analysis: relationship to Ishak staging and elasticity by shear-wave elastography. J Dig Dis 16:217–227. 10.1111/1751-2980.12231 [DOI] [PubMed] [Google Scholar]
  • 54.Kheiri F, Rahnamayan S, Makrehchi M, Asilian Bidgoli A (2025) Investigation on potential bias factors in histopathology datasets. Sci Rep 15:11349. 10.1038/s41598-025-89210-x [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 55.Franchet C, Schwob R, Bataillon G, Syrykh C, Péricart S, Frenois FX, Penault-Llorca F, Lacroix-Triki M, Arnould L, Lemonnier J, Alliot JM, Filleron T, Brousset P (2024) Bias reduction using combined stain normalization and augmentation for AI-based classification of histological images. Comput Biol Med 171:108130. 10.1016/j.compbiomed.2024.108130 [DOI] [PubMed] [Google Scholar]
  • 56.Naik SN, Forlano R, Manousou P, Goldin R, Angelini ED (2023) Fibrosis severity scoring on Sirius red histology with multiple-instance deep learning. Biol Imaging 3:e17. 10.1017/S2633903X23000144 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 57.Ko SM, Shin J-I, Hong Y, Kim H, Sohn I, Lee J-Y, Han H-J, Jeong DS, Lee Y, Son W-C (2025) Deep learning-based method for grading histopathological liver fibrosis in rodent models of metabolic dysfunction-associated steatohepatitis. Front Med 12:1629036. 10.3389/fmed.2025.1629036 [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Data Availability Statement

The datasets used and/or analyzed during the current study are available from the corresponding author upon reasonable request.


Articles from Toxicological Research are provided here courtesy of Korean Society of Toxicology

RESOURCES