Skip to main content
Poultry Science logoLink to Poultry Science
. 2026 Feb 26;105(5):106721. doi: 10.1016/j.psj.2026.106721

KNLR: A heterogeneous ensemble learner for predicting Foie gras weight grade in mule ducks (Anas platyrhynchos × Cairina moschata)

Jia-Cheng Li a, Ichraf Mabrouk a, Qiu-Yuan Liu a, Sheng-Yi Li a, Xiao-Ming Ma a, Yu-Pu Song a, Jing-Yun Ma a, Yu-Xuan Zhou a, Jia-Hua Shao a, Xin-Yue Li a, Jing-Bo Wang a, Gui-Zhen Xue a, Hong-Xiao Pan a, Jing Xu a, Guo-Qing Hua a, Jia-Lin Zhang a, Jun Zhang d, Wei Min d, Fu-Jun Zhang e, Ying-Wei Ma d, Hao Shi d, Yong-Feng Sun a,b,c,⁎
PMCID: PMC12970397  PMID: 41780494

Abstract

Ante-mortem prediction of foie gras weight grade remains an unsolved challenge in commercial duck production. We developed KNLR, a novel heterogeneous ensemble learner that accurately predicts foie gras weight classification in mule ducks using pre-overfeeding morphometric measurements and post-overfeeding live weight, enabling producers to optimize feeding strategies and improve grading consistency.

KNLR integrates Heterogeneous Ensemble Feature Selection (HEFS) with Weighted Area Under Curve Evaluation (WAUCE) to enhance predictive robustness. Comparative evaluation with four base learners (LightGBM, Naïve Bayes, Random Forest, and K-Nearest Neighbors) indicated that KNLR achieved the best overall performance across multiple machine-learning and statistical metrics. Using six features, KNLR achieved the highest precision (0.6425 ± 0.0869), significantly outperforming all base learners. Feature importance analysis indicated that overfeeding liver weight, breast depth, and body slope length were the most important predictors of foie gras grade. The proposed heterogeneous ensemble model may allow early identification of mule ducks with high-quality livers, providing a basis for precision feeding strategies aimed at optimizing feed efficiency and foie gras quality. By supporting grade-specific feeding management during the overfeeding period, KNLR offers a data-driven approach for breeding enterprises to potentially reduce production costs and improve economic returns through more accurate liver grade prediction.

Keywords: Heterogeneous ensemble learning, Foie gras, Feature selection, Feature importance, Mule duck

Introduction

Foie gras production exploits the physiological capacity of waterfowl to store energy through hepatic lipogenesis (Pilo and George, 1983). This delicacy is produced by short-term overfeeding of mature ducks with high-energy diets, inducing controlled hepatic steatosis that yields livers with exceptional nutritional value—many lipids such as free fatty acids, phospholipids and so on (Trehiou et al., 2025). Modern foie gras production primarily uses mule ducks (Anas platyrhynchos × Cairina moschata), produced by crossing Muscovy drakes (Cairina moschata) with Pekin ducks (Anas platyrhynchos) (Baéza, 2005). This cross is characterized by high hepatic lipogenic efficiency and fast growth rate (Baéza, 2005; Flament et al., 2012).

Despite China's steady expansion in foie gras production since the early 1980s, substantial quality inconsistencies persist. Average liver weights stagnate at approximately 600 g, reflecting suboptimal production performance and considerable individual differences (Chen and Chen., 2007; Zhang et al., 2015). This instability undermines both economic benefit and product standardization, creating an urgent need for improved selection and management strategies.

Current methodologies for assessing foie gras weight present critical limitations. Direct post-overfeeding morphometric measurements risk hepatic rupture due to extreme liver enlargement, violating fundamental animal welfare principles. Although computed tomography is noninvasive, its high-cost limits practical application. To date, no validated approach is available for ante-mortem prediction of foie gras weight grade, and breeding enterprises therefore depend on post-slaughter assessment, precluding timely management interventions.

The development of predictive learners using pre-overfeeding morphometric data combined with post-overfeeding live weight offers transformative potential. In this context, the development of such an approach would enable precision feeding protocols tailored to individual liver production capacity, minimize handling stress through targeted selection, reduce unnecessary overfeeding, and maintain animal welfare standards while optimizing production efficiency. Addressing this methodological gap is therefore essential for improving both the economic sustainability and ethical standards of foie gras production.

Big data analytics have revolutionized livestock production through their capacity to extract actionable insights from heterogeneous structured and unstructured datasets (Andreu-Perez et al., 2015; Arshad et al., 2024; Deng et al., 2025; García-Vázquez, 2024). In recent years, machine learning has emerged as a key computational framework for phenotypic prediction in major livestock species, including swine, cattle, and sheep, supporting the development of data-driven selection strategies tailored to specific production traits(Ali et al., 2015; Alonso et al., 2015; Coşkun et al., 2023; Csóka et al., 2025; Hamadani et al., 2022; Huma and Iqbal, 2019; Mota et al., 2021; Tırınk et al., 2023; Zhou et al., 2022; Zhu et al., 2025).

However, most current studies on machine learning in livestock phenotype prediction employ conventional feature selection techniques or bypass feature engineering. This is particularly evident for complex production traits such as foie gras weight classification.

Ensemble learning architectures have emerged as powerful solutions to the limitations of single-algorithm approaches, systematically combining multiple learners to enhance predictive performance. These frameworks comprise two principal paradigms: homogeneous ensembles, which aggregate predictions from multiple instances of a single algorithm, and heterogeneous ensembles, which integrate diverse base learners through sophisticated fusion strategies (Hsu and Srivastava, 2009; Tao et al., 2019; Yu and Xie, 2019; Zefrehi and Altınçay, 2020).

Heterogeneous ensemble methodologies have been thoroughly verified for diverse medical applications including cancer prognosis prediction, coronary heart disease diagnosis and clinical decision support systems (Bashir et al., 2015; Feng et al., 2022; Thongkam et al., 2008; Velusamy and Ramasamy, 2021). This combination of different model architectures usually improves generalization performance and robustness to overfitting (Hsu and Srivastava, 2009; Yang et al., 2010; Sesmero et al., 2015; Zefrehi and Altınçay, 2020). Nonetheless, heterogeneous ensemble applications remain glaringly underexplored in livestock phenotype prediction. Although these modern techniques are proven better in complex classification tasks, they have not been applied to agricultural sectors systematically.

We present ‘KNLR’, a heterogenous ensemble learning model to develop ante-mortem learner of weight grades for mule ducks by production traits of foie gras. KNLR combines four complementary base learners that capture different morphological and weight-related patterns. The ensemble architecture makes use of Heterogeneous Ensemble Feature Selection (HEFS) to select optimal predictor subsets while maintaining algorithmic diversity. Then, it uses a Weighted Area Under Curve Ensemble (WAUCE) to assign dynamic weights to the contributions of base learners according to their classification performance for each foie gras grade category.

We conducted an evaluation and comparative analysis of the machine learning and statistical metrics for KNLR and the base learners. We performed feature importance analysis on base learners to find the morphometric and weight variables that were most highly associated with foie gras grade outcomes. Through the provision of a strong predictive model and significant biologically relevant predictors, this study provides duck producers with real tools for precision selection, targeted feeding and effective quality control thereby enhancing production efficiency and economic impact and a methodology for other research in precision poultry production.

Materials and methods

Ethics statement

All experimental protocols involving the use of animals were conducted according to the guidance approved by the Animal Welfare and Ethics Committee, Jilin Agricultural University (Changchun, Jilin Province, China; approval number: 2025 06 09 001).

Data collection and preprocessing

A total of 646 male mule ducks aged 6 weeks were subjected for a 3-week restricted feeding trial. Ducks were supplied by Jilin Zhongyi Food Technology Co., Ltd and Jilin Zhengfang Animal Husbandry Co., Ltd. Ducks had ad libitum access to feed for two hours during the preparation stage. To induce appetite and maximize intake capacity before overfeeding or for developing normal feeding strategies, this is the practice in mule duck production (Guémené and Guy, 2004).

At 9 weeks of age, eight morphometric traits were measured prior to overfeeding: LW (live weight before overfeeding, kg), BA (breast angle, measured as the angle between bilateral pectoral muscles at the keel in supine position, °), BSL (body slope length, distance from the anterior superior articulation of the clavicle to the ipsilateral ischial tuberosity, cm), FBL (keel length, distance from the anterior to posterior extremity of the sternal keel, cm), BW (breast width, distance between bilateral shoulder joints, mm), BD (breast depth, perpendicular distance from the first thoracic vertebra to the anterior margin of the keel, mm), SL (shank length, linear distance from the superior tarsal joint to the interdigital space between the third and fourth digits, cm), and SG (shank girth, circumference at mid-shank, cm) (National Standardization Administration, 2024).

After morphometric measurements, ducks underwent a 3-week overfeeding period. They were housed individually and force-fed four times daily at 6 h intervals using pneumatic gavage equipment. The initial feed amounts at the start of each session were between 100 and 120 g and were gradually increased to 250–300 g at the end of the overfeeding period. The force-feeding mixture contained 60% water, 38% maize. At the end of overfeeding, the live weight at the end of overfeeding (OLW, kg) was measured and added as a predictor variable.

The ducks were afterwards killed and the livers removed to determine the weight. Via longitudinal incision the carcass was opened, liver was separated from other viscera and liver weight (g) was recorded (National Standardization Administration, 2024). Foie gras grade (FGG) was assigned on the basis of the following criteria.

  • •

    Grade A: liver weight ≥ 600 g

  • •

    Grade B: liver weight < 600 g

All measurements were recorded to four decimal places. Missing values were imputed using mean substitution. Data were compiled using Excel 2021 (Microsoft Corporation, Redmond, WA, USA). Descriptive statistics and subsequent machine learning analyses were performed using R software (version 4.5.1; R Foundation for Statistical Computing, Vienna, Austria).

Heterogeneous ensemble learner architecture

KNLR, a heterogeneous ensemble learner, incorporated three integrated modules: heterogeneous ensemble feature selection (HEFS), base learner construction, and weighted area under curve ensemble (WAUCE) integration. HEFS implemented five feature selection algorithms in the modeling workflow and used majority voting to identify the most informative features. These features selected eventually excepted in the training of four heterogeneous base learners. The predictions were achieved by WAUCE which aggregated the predictions of all the base learners via performance-based weighting.

All computational analyses were performed on a workstation equipped with Intel Core Ultra 5 125H processor, RAM 16.0 GB and Windows 11 operating system (Microsoft Corporation, Redmond, WA, USA). Graphical Representation of the whole ensemble learning framework describe in Fig. 1.

Fig. 1.

Fig 1 dummy alt text

Framework of heterogeneous ensemble learner.

Heterogeneous ensemble feature selection

To build the HEFS framework, five feature selection algorithms have been utilized: Information Gain (IG), Correlation (COR), Random Forest importance (rf), Variance (VAR) and Chi-square test (Chi²). The main idea of feature selection algorithm will be integrated by using a majority voting mechanism. The feature selection methods were based on the following computational principles.

1) Information Gain (IG): Features were transformed into categorical variables through quartile-based discretization, after which IG values were computed according to Eq. (1).

IG(Y,X)=H(Y)−H(Y|X) (1)

where H(Y) represents the entropy of the target variable and H(Y|X) denotes the conditional entropy given feature X. Features with IG > 0.1 were retained for subsequent analyses.

2) Correlation (COR): Pearson correlation coefficients were computed between each feature and the target variable, and features with absolute correlation coefficients exceeding 0.1 (|r| > 0.1) were selected for learner training.

3) Random Forest (rf): An rf model comprising 100 decision trees was trained using 5-fold cross-validation. Feature importance scores were quantified as the mean Gini impurity reduction across all trees. For a single split based on feature Xj, the Gini impurity reduction was calculated according to Eq. (2):

ΔGini=Gini(Dp)−(NLNPGini(DL)+NRNPGini(DR)) (2)

The cumulative Gini impurity reduction for feature Xj in tree k was computed as shown in Eq. (3):

Ik(Xj)=∑t∈Tk:split(t)=XjΔGinik,t(Xj) (3)

The mean feature importance across all K trees (K = 100) was determined using Eq. (4):

Importance(Xj)=1K∑k=1KIk(Xj) (4)

where Dp, DL, and DR represent the sample sets at the parent node, left child node, and right child node, respectively; Np, NL, and NR denote their corresponding sample sizes; Xj indicates an individual feature; and Tk represents the set of all internal nodes across all trees in the forest. Features with importance scores exceeding 0.1 were retained for subsequent analyses.

4) Variance (VAR): Calculation of feature variances on all predictor variables was performed and the 10th percentile of variance distribution was selected. Features whose variance values did not exceed the cutoff were removed to eliminate low-variance predictors.

5) Chi-square Test (Chi²): Continuous features were transformed into binary features through thresholding using the median. Subsequently, Chi-square analyses were carried out to evaluate the relationship of every binarized feature with the target variable. Learners were trained on features that show a statistically relevant association (p < 0.05).

A majority voting mechanism was employed to synthesize feature selection results. For each candidate feature, a binary vote was recorded from each of the five selection methods, creating a voting matrix Vm×n where m = 5 (methods) and n = 9 (features). Matrix entries were defined as:

Vij={1,ifmethodiselectsfeaturej0,Otherwise (5)

The optimal feature set used to train each learner was created by keeping features approved by three or more methods (majority threshold: ≥ 60%).

Base learner construction

To build well-structured ensemble learning systems, we used four classification algorithms representing different learning paradigms and hypotheses spaces as base learners that is, K-Nearest Neighbors (KNN, a distance-based geometric model), Naive Bayes (NB, a probabilistic model), Random Forest (RF, a bagging ensemble), and Light Gradient Boosting Machine (LightGBM, a boosting ensemble). With diverse architectures, complementary error patterns, and better generalization performance.

To control overfitting, hyperparameters were set to optimize the model performance. The k parameter for KNN was varied according to the number of samples considered. Smaller values were used for reduced datasets while larger values were used for expanded datasets. This was done to balance between the bias and variance. NB utilized weakly informative priors in order to diminish assumptions related to the distributions of each feature. RF was configured with 300 trees and a minimum leaf sample size of 3, providing a balance between model complexity and overfitting risk. LightGBM was trained at learning rate 0.05 with a maximum leaf nodes of 31, feature fraction 0.9, bagging fraction 0.8 and boosting iteration 100 to optimize gradient descent while maintaining generalization.

Weighted area under curve ensemble

WAUCE is a novel ensemble integration strategy used to aggregate the outputs of base learners. The mechanism assigns weights dynamically to each learner based on its predictive performance. The final ensemble prediction is a weighted average of base learners’ predictions. This performance-adaptive weighting method assigns greater importance to high-performing peers, whilst down-weighting the contributions of weak learners.

For the ensemble comprising M base learners {h1,h2,…,hm}, individual AUC values {AUC1,AUC2,…,AUCm} were computed on the training dataset. Base learner weights were derived through AUC normalization as specified in Eq. (6):

wi=AUCi∑i=1MAUCj (6)

The WAUCE ensemble prediction was determined by comparing weighted posterior probability sums across classes, as formalized in Eq. (7):

PWAUCE(X)={1,∑i=1Mwi*pi1(x)>∑i=1Mwi*pi2(x)2,Otherwise (7)

where pi1(x) and pi2(x) represent the posterior probabilities assigned by base learner i to Grade A and Grade B, respectively, with the constraint pi1(x)+pi2(x)=1. The sample is assigned to the grade that has the highest weighted posterior probability sum.Where high performing base learners, will have greater influence over classification as compared to low performing ones.

Learner performance evaluation

A performance analysis comparing the four base learners and KNLR ensemble was done. The learner evaluation executed 10 repetitions of 10-fold cross-validation utilizing stratified sampling in order to guarantee that class proportions within each fold mirror the entire set-up appropriately. This development assist in reducing the evaluation bias. The classification performance was quantified using standard machine learning metrics such as accuracy, precision, recall, f1-score, ROC curve and the area under the ROC curve (AUC) index at the end of testing and was reported as mean ± SD across all the 100 validations. For reveals mean performance and SD of the classifiers, they were rounded to four decimal places throughout the text. The evaluation metrics were calculated as follows:

Accuracy=TP+TNTP+TN+FP+FN (8)
Precision=TPTP+FP (9)
Recall=TPTP+FN (10)
F1−score=2*Precision*RecallPrecision+Recall (11)
AUC=∫01TPR(FPR)d(FPR) (12)
FPR=FPFP+TN (13)

The results of classification were measured using conventional terminology from the confusion matrix. This was done with reference to the binary foie gras grading system. Grade A was designated as the positive class while Grade B was taken as the negative class.

  • •

    True Positive (TP): Grade A livers correctly predicted as Grade A

  • •

    True Negative (TN): Grade B livers correctly predicted as Grade B

  • •

    False Positive (FP): Grade B livers misclassified as Grade A

  • •

    False Negative (FN): Grade A livers misclassified as Grade B

The True Positive Rate a.k.a recall (TPR) and False Positive Rate (FPR) are calculated using the Eqs. (10) and (13), respectively.

Paired comparisons were used for the statistical validation of the KNLR ensemble. To evaluate performance differences between KNLR and each base learner, we use the p-values calculated from the paired t-tests over the 100 cross-validation iterations. Kendall’s Tau, τ, for ordinal agreement between learner predictions and true foie gras grades is also computed. It is a distribution free measure of how predictions agree with true classes.

Feature importance analysis

Algorithm-specific feature importance quantification methods were employed for each base learner to ensure compatibility with their underlying computational mechanisms.

KNN. Feature importance evaluated via chi-square test statistics explores the strength of association of a single predictor for classifying foie gras grades.

NB. The significance of the NB indicates the varying power of the distributions of each feature associated with all grades across each category. For each feature, the absolute difference between the class-conditional means was scaled by the pooled standard deviation.

LightGBM. 5-fold Cross Validation was run to train the models. The hyperparmeters were the same as those for LightGBM in Base Learner Construction. The importance scores of features were derived during model training using gain values, which represent the total improvement in split quality that can be attributed to each feature at all boosting stages.

RF. The models were trained on 5-fold cross-validation with hyperparameters as specified for the Base Learner Construction. As documented in Eqs. (2)-4 of the RF component of Heterogeneous Ensemble Feature Selection, feature importance corresponded to the mean Gini impurity reduction.

To ensure stability and robustness, feature importance of the four base learners was evaluated over 30 independent iterations. Every subsequent time, the importance score normalized in the range of [0, 1] for further runs. The 30 normalized vectors were averaged for each learner to yield stable, learner-specific feature importance rankings, thereby mitigating stochastic variability associated with individual model training.

The feature importance scores were averaged across the four algorithms and then used to compute a mean importance score across the base learners in order to achieve a consensus ranking. This averaging procedure, first across iterations within each learner and then across learners for each feature, generated a full importance ranking that did justice to the varying perspectives of all four classifiers. The means of all feature importance values were reported, rounded to 4 decimal places.

Results

Descriptive statistics of predictor variables

Table 1. contains descriptive statistics for all predictor variables. For most features, similar mean and median values suggest the data were approximately symmetrically distributed (Jia et al., 2007). The coefficients of variation (CV) were low ranging from 4.24% to 9.78% (Jia et al., 2007). Similarly, the standard errors (SE) ranged from 0.0133 to 0.4146 The distribution of foie gras grades included 38.08% of Grade A (n = 246) and 61.92% of Grade B (n = 400).

Table 1.

Statistical description of the features used in this study.

Feature Name Min Max Mean Median SD SE CV(%)
LW (kg) 2.3300 4.7300 3.4556 3.4556 0.3378 0.0133 9.7756
BA 100.2000 168.8000 111.1639 111.1639 6.2004 0.2440 5.5777
BSL (cm) 21.2000 31.5000 26.9023 26.9023 1.1409 0.0449 4.2408
FBL (cm) 14.0000 23.0000 17.7730 17.7730 1.1143 0.0438 6.2697
SL (cm) 6.0000 9.0000 7.4273 7.4273 0.4298 0.0169 5.7865
SG (cm) 4.3000 8.5000 5.2375 5.2375 0.3296 0.0130 6.2935
BD (mm) 100.0900 181.2400 119.3547 117.6100 10.5371 0.4146 8.8284
BW (mm) 100.2500 191.7400 111.8365 111.8365 8.3445 0.3283 7.4614
OLW (kg) 4.5198 7.4350 6.2607 6.2607 0.3891 0.0153 6.2143

SD: standard deviation, SE: standard error, CV: coefficient of variation.

Heterogeneous ensemble feature selection

To improve a model's robustness and generalization ability, HEFS was adopted to choose the best predictor subset. The ensemble strategy has utilized five different algorithms of features’ selection, where the final features selected in the majority voting. Fig. 2. presents the voting results of each candidate feature across all the five algorithms.

Fig. 2.

Fig 2 dummy alt text

Voting selection of five feature selection methods.

The majority voting criterion (vote threshold ≥ 3) enabled the retention of 6 features for foie gras grade classification, namely BSL, FBL, SG, BD, BW, OLW. At least three of the five selection algorithms have endorsed these features, suggesting they are useful predictors which are consistently identified in different statistical and machine learning frameworks.

Feature selection robustness was assessed by repeating the HEFS procedure for 30 independent runs. Fig. 3. shows the selection frequency of each feature over those iterations, displaying the most chosen predictors among them. The full voting records for each run are given in the Supplementary Materials.

Fig. 3.

Fig 3 dummy alt text

30 Rounds of heterogeneous feature selection results.

The initial analysis resulted in the identification of six features (BSL, FBL, SG, BD, BW, OLW), subsequently further attempts frequently reproduced the outcome revealing the reproducibility of the HEFS ensemble approach. The framework for feature selection proved to be robust through independent runs.

Learner performance evaluation

As demonstrated in Table 2. and Fig. 4., we examined the performance of all five learners through 10 repetitions of 10-fold cross-validation. KNLR achieved a precision of 0.6425 ± 0.0869, representing a 2.16% improvement over RF (0.6209 ± 0.0791), the best-performing base learner. In terms of overall accuracy, KNLR (0.7080 ± 0.0504) ranked second among all learners, exceeded only by LightGBM (0.7105 ± 0.0531).

Table 2.

The machine learning indicators of six learners after ten rounds of ten-fold cross-validation.

Learner Accuracy Precision Recall F1 Score AUC
KNN 0.6407 ± 0.0578 0.5398 ± 0.1065 0.4143 ± 0.1016 0.4644 ± 0.0943 0.6582 ± 0.0629
NB 0.6349 ± 0.0538 0.5355 ± 0.1184 0.3371 ± 0.1007 0.4088 ± 0.1005 0.6650 ± 0.0802
LightGBM 0.7105±0.0531 0.6194 ± 0.0973 0.6413±0.0973 0.6266±0.0700 0.7889 ± 0.0532
RF 0.7023 ± 0.0467 0.6209 ± 0.0791 0.5765 ± 0.1078 0.5928 ± 0.0755 0.7919±0.0498
KNLR 0.7080 ± 0.0504 0.6425±0.0869 0.5310 ± 0.1117 0.5768 ± 0.0893 0.7838 ± 0.0529

Fig. 4.

Fig 4 dummy alt text

The machine learning index radar chart of different learners.

For the AUC, KNLR demonstrated performance (0.7838 ± 0.0529) comparable to RF (0.7919 ± 0.0498) and LightGBM (0.7889 ± 0.0532), while substantially outperforming KNN (0.6582 ± 0.0629) and NB (0.6650 ± 0.0802). In Fig. 5., we present the aggregated ROC curves corresponding to the 100 iterations of the validation process. The color-coding is as follows KNN-blue, LightGBM-orange, NB-red, RF-purple and KNLR-green. Supplementary Figures S1.- S5. show individual ROC curves for each learner across all cross-validation folds.

Fig. 5.

Fig 5 dummy alt text

ROC curves of different learners.

The first stage involved the use of paired t-tests to evaluate how KNLR performed with respect to the four base learners based on five different performance metrics. As indicated in Table 3., KNLR exhibited superior performance to each of the base learners in respect of precision, recall, F1-score and AUC (p < 0.05). KNLR achieved notable enhancements in accuracy when compared to three of the four base learners, with varying differences in magnitude.

Table 3.

The p-value comparison between KNLR and base learner on machine learning metrics.

Learner p-value (KNLR VS base learner)
Accuracy Precision Recall F1 Score AUC
KNLR - - - - -
KNN 0.0000 0.0000 0.0000 0.0000 0.0000
NB 0.0000 0.0000 0.0000 0.0000 0.0000
LightGBM 0.4895 0.0010 0.0000 0.0000 0.0010
RF 0.0326 0.0008 0.0000 0.0017 0.0010

p-value < 0.001: extremely significant statistical differences; 0.001 ≤ p-value < 0.01: significant statistical differences; 0.01 ≤ p-value < 0.05: statistical differences; p-value ≥ 0.05: no statistical differences.

Kendall's Tau rank correlation coefficient (τ) quantifies the concordance between learner predictions, providing insight into learner diversity and independence (Schaeffer and Levitt, 1956). Fig. 6. presents the correlation heatmap based on predictions across all 100 cross-validation iterations, with p-values indicating statistical significance. All pairwise τ values achieved statistical significance (p < 0.001). The τ values of the base learners are between 0.343 to 0.808. Further, the predictive diversity is substantial.

Fig. 6.

Fig 6 dummy alt text

Heatmap of Kendall 's Tau rank correlation coefficient (τ) based on ten-round ten-fold cross validation. ⁎⁎⁎: p-value < 0.001:; ⁎⁎: p-value < 0.01; *: p- value < 0.05.

The τ values between KNLR and its individual base learners show a clear trend from moderate to strong. This trend shows how WAUCE modifies the contribution of each learner in the ensemble dynamically: the stronger learner makes a more significant contribution to the classification while the weaker learner less. This performance-based weighting method is designed so that KNLR benefits from its best-performing assets.

Feature importance analysis

Feature importance analysis was performed over 30 independent iterations for each of the four base learners to identify the morphometric and weight variables most strongly associated with foie gras grade classification. Table 4. summarizes the overall importance scores, while Fig. 7. visualizes the same. Following the color scheme: KNN (blue), LightGBM (orange), NB (red), and RF (purple). The Supplementary Figures S6.- S9. provide individual bar chart showing the importance of each base learner. The feature importance is ranked in descending order as follows: OLW, BD, BSL, FBL, BW, SG.

Table 4.

The feature importance score of four base learners.

Learner OLW BD BSL FBL BW SG
KNN 1.0000 ± 0.0000 0.0178 ± 0.0000 0.2099 ± 0.0000 0.1784 ± 0.0000 0.1308 ± 0.0000 0.0526 ± 0.0000
NB 1.0000 ± 0.0000 0.5025 ± 0.0000 0.4750 ± 0.0000 0.4696 ± 0.0000 0.0772 ± 0.0000 0.3159 ± 0.0000
LightGBM 1.0000 ± 0.0000 0.2766 ± 0.0107 0.1062 ± 0.0051 0.1028 ± 0.0045 0.2454 ± 0.0071 0.1163 ± 0.0061
RF 1.0000 ± 0.0000 0.4738 ± 0.0055 0.2565 ± 0.0034 0.2670 ± 0.0043 0.4167 ± 0.0055 0.2503 ± 0.0036
Average 1.0000 ± 0.0000 0.3177 ± 0.0060 0.2619 ± 0.0031 0.2545 ± 0.0031 0.2175 ± 0.0045 0.1838 ± 0.0035

Fig. 7.

Fig 7 dummy alt text

The feature importance bar graph of four base learners.

Discussion

Predictive performance and evaluation of learners

In this work, we proposed a new heterogeneous ensemble learner KNLR for ante-mortem prediction of foie gras weight grades in mule ducks that were computed from pre-overfeeding morphometric measurements and post-overfeeding live weight via HEFS and WAUCE respectively. The performances of the above model were evaluated in detailed with those of the 4 base learners, using 6 metrics: accuracy, precision, recall, F1-score, AUC, and ROC curves.

The choice of accuracy as the main criterion of choice reflects the fundamental economic realities of commercial foie gras production. Grade A livers receive much higher market premiums but they only constitute 38.08% of the production output in our dataset which is a characteristic minority class distribution. The resulting imbalance in classes creates a need for reliable predictions as false positives, meaning a Grade B liver being classified as Grade A, would lead to a loss of profit due to downgrading the product. This occurs when the expectation does not match the quality of the product and results in loss of brand reputation. The fraction of Grade A livers that are correctly predicted is focused on the efficient production of foie gras. When KNLR labels a liver as Grade A, it is done with great precision, thereby ensuring that all resources and management efforts are focused on truly top-grade livers. This prevents losses that may otherwise arise from a wrong grade.

The different metrics help give context to the learner evaluation Accuracy reflects the overall correctness of the two grades while recall is the number of Grade A livers which the learner correctly identifies. This shows how sensitive the learner is to the minority class (Zhou, 2021). F1-score harmonizes precision and recall, preventing optimization of one at the expense of the other (Zhou, 2021). Ultimately, ROC curves and AUC assess the discriminative ability under different classification thresholds and characterizes the primary trade-off between true positive and false positive rates of binary classification (Zhou, 2021). These metrics combined indicate that possessing a high performance requires having high precision (commercial reliability), acceptable recall (capture rate of premium product), and strong AUC (good class separation). KNLR certainly has this capability.

KNLR's performance in the identification of Grade A liver was superior with a precision of 0.6425 ± 0.0869 which is significantly higher than all constituent base learners. Even if this precision is low in absolute terms, it has significant practical value given appropriate context. Relative to random classification, whose baseline precision is 38.08% (which is the prevalence of Grade A), KNLR improves relative identification accuracy by 68.72%. This increase in performance yields economic benefits. Since, Grade A livers earn 40% more profit than Grade B livers, despite the 36% misclassification rate, the usage of KNLR-informed selection strategies would enhance the overall production profit by around (64.35% −38.08%) × 40% ≈ 10.5%. Beyond immediate classification benefits, KNLR enables proactive management interventions: producers can adjust overfeeding duration, modify dietary formulations, or implement grade-specific feeding protocols based on predicted liver grade potential, optimizing feed conversion efficiency while maximizing premium product yield. The combined capabilities of KNLR make it a precision decision-support tool that helps improve product quality control, stratified marketing and production economics – using accessible and non-invasive phenotyping to address the needs of an industry.

This study adopted a different approach compared to recent studies whereby imaging technologies were used to predict poultry traits. For example, Csóka et al. (2025) estimated breast and leg muscle weights in chicken using computed tomography (CT), and Vali et al. (2023) used CT imaging to measure spleen size in broilers. Xu et al., (2018) utilized CT scan technology to evaluate liver fat content in 22 Landaise geese through overfeeding. However, a critical implementation challenge exists to applying CT-based approaches in commercial-scale foie gras production. Because of the high cost per scan, it is not economically viable to screen large flocks. Furthermore, the restrained immobilisation of the birds required for imaging after post-overfeeding raises serious animal welfare issues: the liver is greatly enlarged and hyperplastic, making it susceptible to traumatic rupture during manipulation and positioning prior to the CT acquisition. In comparison, our morphometric and weight-based approach only employs routine husbandry measurements that do not rely on special equipment or liver-threatening restraint procedures. It also fits well within the budget and operational practices of commercial duck production systems. Thus, machine learning on simple morphometric data successfully implements a strategic paradigm: reaching acceptably predictive performance with accessible phenotypes may be more impactful to industry than superior accuracy at prohibitive cost and/or invasiveness.

Heterogeneous ensemble feature selection

Feature selection is essential in machine learning and can be divided into three paradigms; filter methods assess features based on statistical independence, wrapper methods assess features jointly with a specific classifier, and embedded methods perform feature selection as part of the training of a model (Zhou, 2016). All three works aim at tackling high-dimensional data challenges through the exclusion of repetitive features, identification of optimal subsets, reduction of model complexity, and enhancement of predictive performance. However, they all differ in approach and assumptions (Cui et al., 2018; Feng et al., 2022; Huang et al., 2019; Wang et al., 2020). Even though there have been many advancements, traditional feature selection methods usually rely on one algorithm, therefore being biased and limited to only that algorithm.

HEFS tackles this important restriction with the combination of five theoretically distinct feature selection techniques which are IG, COR, rf, VAR, and Chi² through majority voting ensemble integration. HEFS also include complementary perspectives of these methods on relevance. The variety of algorithms employed in this research offers a multi-faceted evaluation of the predictive value of features. Therefore, IG quantifies the reduction of entropy that a certain feature brought about in the foie gras grade through its nonlinear dependencies. Additionally, the COR directly assesses linear dependencies and is particularly effective for continuous morphometric traits as it is intuitively interpretable. Moreover, rf importance indicates the predictive power of a feature due to its measurement in the recursive partitioning algorithm. Furthermore, VAR filtering eliminates features with low variance which do not account for the theoretical variance of that feature. Finally, Chi² assesses the statistical significance of associations between discrete grade categories and binarized features. This formal hypothesis testing is coupled with variance analysis to provide a statistically rigorous foundation for retaining features (Benaddi et al., 2025; Odhiambo Omuya et al., 2021; Zou et al., 2003).

The integration mechanism of the HEFS represents a new methodological innovation. HEFS enshrines a conservative selection criterion through majority endorsement requiring at least three of five methods for a feature to be selected. This guarantees robust and consistently validated predictors being selected, opposing selection of method-specific artifacts. These selections are stable across 30 independent run simulations and cross-validation procedures (Fig. 3.), which offers empirical evidence for their reliability. Significantly, the ensemble approach enhances predictive accuracy and interpretability. Indeed, when a feature receives support from statistical tests (Chi², VAR), correlation analysis (COR), information-theoretic measures (IG), and machine learning algorithms (RF), its biological relevance is reinforced from several perspectives. The overlap of evidence will give more confidence about feature importance ranking than what single method selection can provide.

HEFS can be used for more than this study. The presented framework is a versatile tool for optimizing features in agricultural machine learning problems. It is useful when the dataset is moderate in size and a simpler model is preferable. Additionally, it is important to be able to interpret the model since practitioners are less likely to use black-box methods. Lastly, predictive performance should be validated across a range of methods in order to demonstrate robustness. The consistency of HEFS-selected features and their convergent validation point to broader application possibilities in precision livestock farming, where accurately predicting phenotypes from limited, high-dimensional data is a particular challenge. The HEFS approach advances the methodological robustness and practical reliability of data-driven animal production systems through the transformation of feature selection from a single-algorithm decision to a democratic agreement.

Base learner construction

Ensemble learning is described as taking many weak learners and aggregating them to form one strong learner that is better able at generalizing. The use of taking many weak learners and aggregating them was first done by Dasarathy (1979). There are three main paradigms of modern ensemble methods that are founded on relationships of base learners. Bagging aims to reduce variance by training many models independently in parallel (Breiman, 1996). Algorithms that are designed to boost are those which look to minimize bias by sequentially correcting the errors made by previous models (Schapire, 2003). Stacking algorithms take advantage of the different strengths of heterogeneous models by using the meta-learning approach (Sigletos et al., 2005). KNLR realises the Stacking paradigm. It utilizes the WAUCE approach to bring four base learners which are strategically different from each other together such that their powers are combined while their weaknesses are circumvented.

The architecture of KNN, NB, LightGBM, and RF was specifically devised to ensure maximum algorithmic diversity through their respective bias-variance levels. Table 5. provides the introduction and function of each basic learner. The aim of using heterogeneous base learners is to ensure that their errors have little in common. For instance: what distance-based methods fail to see, a probabilistic approach might pick up on. And what linear models fail to approximate, tree-based learners may be constructed to deal with. WAUCE achieves this diversity of learners by assigning weights based on their respective performance on validation data. In other words, a merit-based aggregation strategy is used where learners performing better on the foie gras grading problem will contribute more to the ensembles’ final predictions.

Table 5.

Description, comparison and function of base learners.

Learner Description Strengths Weaknesses Function
KNN (Boateng et al., 2020) Geometric model based on sample similarity prediction Big data effect is obvious, non-parametric model, simple operation The effect of high-dimensional data is weak Supplement local information
NB (Patil and Sherekar, 2013) Probabilistic model based on Bayesian theorem Strong interpretability and is suitable for high-dimensional data Nonlinear relationship capture ability is weak Supplementary linear relationship
LightGBM (Pan et al., 2020) Boosting algorithm based on decision tree Low memory consumption, automatic processing of non-linear relationship Poor interpretability, small data may be over-fitting Supplementary nonlinear relationship
RF (Breiman, 2001) Bagging algorithm based on decision tree The risk of overfitting is low, the feature combination is automatically explored Poor interpretability, slow speed Provide robustness and reliability

The decision not to include certain well-known algorithms in the base learner pool has been taken, not accidentally but for reasons of methodology. The usage of AdaBoost is quite common and widely accepted, nonetheless, it exhibits a prominent sensitivity to outliers and label noise. This is a crucial fact, especially when working with biological measurements which, by nature, will experience recording errors. These errors regularly originate from phenotypic variability and measurement inaccuracy (Feng et al., 2022). The real-world agricultural datasets are often noisy, and AdaBoost re-weighting of misclassified samples can lead to enhanced spurious patterns rather than the true signal. Deep neural networks, while highly effective for extracting patterns from high-dimensional data, were excluded from this study due to practical and interpretability constraints. Because of their high computational requirements, they are not suitable for on-farm application, and the black-box nature of their forecasts limits transparency that is essential for producer trust and regulatory acceptance in agtech. Since SVM is usually in the geometric model, using it together with KNN as a base learner will lead to redundancy within the geometric model category (Bourouba et al., 2025; Mavroforakis and Theodoridis, 2005). This may potentially reduce the structural heterogeneity and robustness of KNLR, rather than enhance it. Therefore, we also did not incorporate SVM as a base learner.

These design choices reflect a principled design philosophy: robustness, efficiency and interpretability rather than merely prediction accuracy. The four chosen base learners offer a combination of diverse inductive biases, computational tractability, and interpretable decision logic, all of which are critical characteristics of a real-world agricultural decision support system. The experimental results of KNLR (precision: 0.6425 ± 0.0869) show that this design strategy can achieve practically useful accuracy but does not have the complexity, data-hungry nature, or opacity that would prevent industrial use. By managing the complexity of our algorithm to match our operational capabilities, we exemplify the design principles for translating machine learning innovations from academic benchmarks to operational agricultural technology.

Overfitting considerations

Based on the grade distribution of foie gras, the class imbalance observed was Grade A: 38.08%, Grade B: 61.92%. This required a strong validation of the generalization performance without overfitting to the dataset. Using arbitrary ratios to partition data into train-test sets, such as 70:30, is quite risky in imbalanced classification scenarios. That's because samples randomly drawn may end up with either the training or test set being heavily weighted towards one class. This inflates performance estimates on validation data that are therefore non-representative of the population. Further, the performance on validation data becomes less important if the model is not able to generalize to the production population. In recent research studies (Chen et al., 2023; Pirompud et al., 2025), this risk was highlighted. To avert this methodological snafu, we employed stratified sampling such that all datasets were constrained to the original 38.08:61.92 class ratio, thus ensuring learner training and testing occurs on representative batches.

To prevent overfitting, besides stratified partitioning, a 10-independent repetition 10-fold cross-validation scheme (that is 100 training-validation cycles in total). In every iteration, the data will be split into ten equal folds, and training is done on nine folds while one fold will be used for validation. Thus each sample will serve as validation data once in each repetition (Zhao et al., 2024). This method imposes stringent separation between the training and test sets within each fold and assesses the stability of performance across multiple subsets of data, benefits that standard holdout validation cannot offer (Arlot and Celisse, 2010).

When performance metrics are reported as mean ± standard deviation over all 100 iterations, it provides a transparent quantification of model stability and addresses generalization issues directly. The KNLR learner achieved low standard deviations in precision (0.0869), accuracy (0.0504) and AUC (0.0529). The values demonstrate consistent performance by KNLR regardless of training-validation split. This shows that KNLR learns patterns instead of memorizing the idiosyncrasies of the training set. Complementary statistical validation through paired t-tests and Kendall's Tau correlation analysis (Table 3., Fig. 6.) further corroborates that observed performance differences reflect genuine algorithmic superiority rather than random variation or overfitting to particular data folds.

The repeated k-fold cross-validation with stratified sampling at each fold, stability assessment, and significance testing show that the performance of KNLR is generalization over unseen data in production. The availability of validating methods from different sources gives strong confirmation that the model will be accurate for new flocks and different production cycles as well as beyond the training dataset of the commercial setting. The KNLR has been validated with a rigorous degree of validation that far exceeds what would normally be expected for agriculture machine learning. As such, the KNLR is a likely candidate for use in practical applications.

Feature importance analysis

Predictive accuracy reflects the ability of a classification model to perform well on a given dataset, but increasing attention is being paid to the mechanistic basis of feature importance and its impact on breeding and management decision-making. Just prediction gives producers the probabilities of outcome, but it does not indicate how to bias selection on certain phenotypic traits or how pre-overfeeding morphology affects the final liver quality. Analysis of feature importance help in eliminating this information gap by showing the contribution of each predictor to grade classification, and thereby clarifying the biological pathways between morphometric traits and the capacity for foie gras production.

In all four base learners OLW was the most important predictor, with scores far exceeding those of any of the pre-overfeeding traits. The study supports the basic principles of physiology as mass of body after overfeeding is a direct reflection of lipid deposition in the liver as the larger the liver, the greater the proportion of total body weight of successfully force-fed waterfowl. The OLW reflects the joint influence of metabolic capacity, feed conversion efficiency, and intrinsic hepatic lipogenic potential and thus represents a composite biomarker of foie gras production performance. The predictive power of the index clearly demonstrates its biological meaning and shows that pre-overfeeding morphometric traits can potentially complement selection decisions made before the costly overfeeding.

Among the five morphometric parameters determined before overfeeding for selecting mouthing index based computation of power function, BD and BSL ranked second and third in importance followed by FBL, BW and SG. This pattern of hierarchy displays the anatomical limitations of hepatic enlargement capacity. The liver of avian. Their bilobed organ lies in the cranioventral abdominal cavity, enveloped by the hepatoperitoneal membrane. The larger right lobe (cardiac-shaped) and the smaller left lobe (rhomboid) connected by an isthmus and coming into contact with the spleen, proventriculus, gizzard, duodenum, jejunum, and reproductive organs (Wu and Xiao, 2023). The thoracoabdominal cavity limits the spatial volume in which the liver can expand during overfeeding (Zhu et al., 2013). BD estimates the area of the body cavity which holds the liver. BSL, records abdominal length over time The framework of the thoracic cavity has complementary dimensions.

The fact that BD and BSL are better predictors of fat accumulation than BW is consistent with the cranioventral orientation of the avian liver and the organ´s expansion pattern during lipid uptake. Over the years, various reports have made claims regarding the nature and mechanisms underlying hepatomegaly accompanied by organ cirrhosis and ascites in special waterfowl. Consequently, BD and BSL more closely determine the potential maximum size of the liver, hence, they are more important in predictions. The anatomical point of view indicates that deep-bodied ducklings with long trunks and pronounced sternal keels should increase foie gras production by enabling better hepatic expansion.

While SG was ranked as the least important amongst the features selected to be retained for predictions, it was consistently selected in all feature selection methods and has some predictive power. This seemingly marginal aspect of bone structure likely affects foie gras production through biomechanics rather than metabolism. Overfeeding results in an increase in body mass that dramatically increases the loading on the hindlimb skeletal structure. Insufficient tibiotarsal robustness may reduce postural stability and mobility while increasing stress responses that negatively affect feed intake and lipid metabolism according to Zhu et al., 2013. Therefore, SG represents the skeletal robustness and load-bearing strength of the species, which indirectly subserves sustained overfeeding tolerance and hepatic lipogenesis.

This mechanistic work will yield practical selections that could extend the utility of KNLR from liver grade to population improvement. To take advantage of vertical space for liver expansion, producers can select ducks with deep-bodied conformation. Trunk length may help increase longitudinal abdominal capacity. Sternal development ought to reflect thoracic volume. Skeleton structure may help support mass gain in overfed stage.

Combining KNLR’s grade predictions with feature importance rank scores, producers can design multi-trait selection indices that optimize current liver production (vary based on OLW through culling after overfeeding) and the future genetic potential (vary based on morphometric through parent stock selection before reproduction). This two-stage selection system selection of breeding candidates based on phenotype before overfeeding and placement of animals in production based on their weight after overfeeding constitutes a precision phenotyping system that allows short-term production quality control and long-term genetic improvement.

Limitations

Despite KNLR demonstrating practical utility to predict foie gras grade, certain methodological limitations should be taken into account and show promising avenues for future research. The morphometric measurements used in this study were obtained manually that is operator-dependent which is difficult to standardize across technicians or production plants. Additionally, the predictor set was mainly derived from pre-overfeeding traits, with OLW being the only trait that encapsulated the state at the lipogenic phase. In the future, it will be necessary to carry out automized phenotyping by integrating computer vision-based systems that are capable of morphometrically acquiring high-throughput samples at standardized conditions that exhibit no inter-observer measurement bias. In addition, it may be possible to expand the feature set to intermediate measures of body weight dynamics, behavioral proxies of feeding efficiency, non-invasive measures of liver enlargement during the course of the overfeeding period. Such measures would enhance the development of trajectory-based classification models that use ‘hidden’ dynamic information not captured at either pre- or post-overfeeding static measures.

Second, rather than performing a feature engineering process specific to that algorithm, the feature selection framework uses a common optimal subset across all base learners. Although HEFS reached a parsimonious six-feature set through ensemble consensus, individual learners may benefit from the use of a predictor that aligns with its algorithmic assumptions. For instance, distance-based approaches such as KNN should concentrate on homogeneously scaled features, while tree-balanced algorithms may take advantage of very informative features, even if they differ in these characteristics. Future work can consider learner-specific feature selection in an ensemble framework, which can boost the performance of individual base learners while maintaining the architectural diversity needed for ensemble integration.

Third, the current learner only uses observed morphometric and weight phenotypes seen, and does not utilize the physiological, biochemical, and genomic information which may contain additional predictive signal. Including oxidative stress markers, such as malondialdehyde levels and superoxide dismutase activity, could further capture individual variability in metabolic resilience to overfeeding. Similarly, adding genetic information, ssSNPs associated with hepatic lipogenesis pathways, would make for a genomics-informed prediction capturing the heritable variation in foie gras production potential. The future of precision selection systems involve combining phenotypic, metabolic and genomic information that add more to mere external morphology by building predictive model from the integration of biological machinery. Nonetheless, this multimodal integration needs to weigh incremental predictive advantages against the operational complexity and cost of commercial situations to acquire molecular phenotypes.

Moreover, the validation study was conducted within a single production system of a facility with male mule ducks of the same age (12-week at slaughter) and by uniform husbandry practices. While this experimental setup achieves consistent data, minimizing the influence of confounding variation, this necessarily confines immediate generalizability across the heterogeneous context of commercial foie gras production.

To confirm its robustness in the face of real-world variability, KNLR needs to be appraised over a wider range of production conditions, including female ducks, reverse Pekin × Muscovy crosses, different overfeeding regimes (duration, diet composition, and frequency of overfeeding), and different sites with varying management practices. This type of validation would clear how much model recalibration or transfer learning strategies are needed for the reliable deployment.

Ultimately, empirical testing is essential for determining whether KNLR’s heterogeneous ensemble architecture can assist with other difficult phenotype prediction problems in animal agriculture. Incorporating some of the other methods, specifically HEFS for a robust feature selection, diversity among the base learners and WAUCE for performance-weighted integration, are also generalizable design principles. However, to predict other hard-to-measure traits such as meat quality traits, disease susceptibility, and reproductive performance will require systematic work. The successful adaptation of KNLR to other agricultural prediction contexts would see it firmly established, not just as a tool for grading foie gras, but as a wider methodological template for precision livestock phenotyping.

While these limitations are significant, they're not grounds for rejecting the product. The performance shown by KNLR based on readily available morphometric measurements demonstrates proof-of-concept of ensemble-based livestock phenotype prediction while the extensions identified automated phenotyping, learner-specific features, multimodal data fusion, multi-environment testing, and cross-trait generalizability provide a structured roadmap for equipping research programs for ever more complex and generalizable precision animal production technologies.

Conclusions

This study designed and validated a heterogeneous ensemble learner called KNLR, an accurate ante-mortem prediction of weight grade in foie gras of mule ducks via combination of morphometric measurements before overfeeding and live weight after overfeeding. The implementation of the two new methodologies namely Heterogeneous Ensemble Feature Selection (HEFS) which is a consensus based methodology where five feature selection methods get together to identify structure robust predictors along with Weighted Area Under Curve Ensemble which is a performance-based methodology adapting four algorithmically diverse base learners significantly improves the classification accuracy as compared to individual models. Through this framework, KNLR produced Grade A liver in which its precision was 0.6425 ± 0.0869. This is well above randomness and shows a reliable discrimination between premium foie gras and standard ones.

The predictive effectiveness shown is linked to production management capacity. KNLR allows an early identification of high potential candidates before the high-cost overfeeding stage. This will empower producers to execute precision feeding strategies whereby expected liver grade will dictate feeding prescriptions. It will allow proper resource allocation towards ducks with higher production potential and set grade-wise husbandry standards in order to effectively convert feeds to other desirable outputs while minimizing production costs. The practical benefits of feature importance analysis mean that once-weighted live weight, breast depth and body slope length are the most influential determinants of foie gras grade, which can inform both immediate culling decisions and longer-term genetic improvement strategies. Therefore, it offers the commercial duck industry a non-intrusive, economical decision-support tool that complements quality assurance, stratified marketing and optimization of production efficiencies.

The KNLR can serve as a generalisation of many other phenotype prediction problems across animal agriculture, instead of just limiting itself to foie gras grading. The design principles of ensemble-driven feature selection for robust predictors, deliberate combination of complementary learning algorithms and performance-weighted aggregation for optimizing the predictive power of the ensemble provide a general concept to build precision phenotyping tools for traits that are difficult or expensive to measure directly on live animals. As the livestock sector adopts increasing data-driven management strategies, methods like KNLR that extract maximum predictive power from available low-cost non-invasive measurements while maintaining interpretability and computational feasibility will be crucial to converting machine learning advances into real-world improvements in animal production systems.

Funding

This work was supported by the Jilin Provincial Department of Science and Technology: Jilin Province Enterprise 'Science and Technology Innovation Specialist' Project (2025); Jilin Provincial Department of Science and Technology: Jilin Province Scientist Studio Project (2025); Jilin Provincial Animal Husbandry Bureau Project (2026) . The authors are grateful for the financial support received for this study.

CRediT authorship contribution statement

Jia-Cheng Li: Writing – review & editing, Writing – original draft, Visualization, Validation, Supervision, Software, Resources, Methodology, Investigation, Formal analysis, Data curation, Conceptualization. Ichraf Mabrouk: Writing – original draft, Supervision, Formal analysis. Qiu-Yuan Liu: Writing – review & editing, Writing – original draft, Formal analysis. Sheng-Yi Li: Writing – review & editing, Writing – original draft, Formal analysis. Xiao-Ming Ma: Writing – original draft, Resources, Investigation, Data curation. Yu-Pu Song: Formal analysis, Data curation, Writing – review & editing. Jing-Yun Ma: Formal analysis, Data curation, Writing – review & editing. Yu-Xuan Zhou: Formal analysis, Data curation, Writing – review & editing. Jia-Hua Shao: Writing – original draft, Writing – review & editing. Xin-Yue Li: Formal analysis. Jing-Bo Wang: Formal analysis. Gui-Zhen Xue: Formal analysis. Hong-Xiao Pan: Formal analysis. Jing Xu: Formal analysis. Guo-Qing Hua: Formal analysis. Jia-Lin Zhang: Formal analysis. Jun Zhang: Resources. Wei Min: Resources. Fu-Jun Zhang: Resources. Ying-Wei Ma: Resources. Hao Shi: Resources. Yong-Feng Sun: Writing – review & editing, Writing – original draft, Supervision, Resources, Project administration, Funding acquisition, Conceptualization.

Disclosures

The authors declare the following financial interests/personal relationships which may be considered as potential competing interests:

Yongfeng Sun reports financial support was provided by Jilin Provincial Department of Science and Technology. Yongfeng Sun reports financial support was provided by Jilin Animal Husbandry Bureau. If there are other authors, they declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper.

Acknowledgments

The authors express their acknowledgment to Jilin Zhongyi Food Technology Co., Ltd and Jilin Zhengfang Animal Husbandry Co., Ltd for providing the mule ducks.

Footnotes

Supplementary material associated with this article can be found, in the online version, at doi:10.1016/j.psj.2026.106721.

Appendix. Supplementary materials

mmc1.zip (293.8KB, zip)

References

  1. Ali M., Eyduran E., Tariq M.M., Tirink C., Abbas F., Bajwa M.A., Baloch M.H., Nizamani A.H., Waheed A., Awan M.A. Comparison of artificial neural network and decision tree algorithms used for predicting live weight at post weaning period from some biometrical characteristics in Harnai sheep. Pak. J. Zool. 2015;47:1579–1585. [Google Scholar]
  2. Alonso J., Villa A., Bahamonde A. Improved estimation of bovine weight trajectories using Support Vector Machine Classification. Comput. Electron. Agric. 2015;110:36–41. [Google Scholar]
  3. Andreu-Perez J., Poon C.C.Y., Merrifield R.D., Wong S.T.C., Yang G.-Z. Big data for health. IEEe J. Biomed. Health Inform. 2015;19:1193–1208. doi: 10.1109/JBHI.2015.2450362. [DOI] [PubMed] [Google Scholar]
  4. Arlot S., Celisse A. A survey of cross-validation procedures for model selection. Stat. Surv. 2010;4:40–79. [Google Scholar]
  5. Arshad M.F., Pietro Burrai G.., Varcasia A., Sini M.F., Ahmed F., Lai G., Polinas M., Antuofermo E., Tamponi C., Cocco R., Corda A., Parpaglia M.L.P. The groundbreaking impact of digitalization and artificial intelligence in sheep farming. Res. Vet. Sci. 2024;170 doi: 10.1016/j.rvsc.2024.105197. [DOI] [PubMed] [Google Scholar]
  6. Baéza E. The fattening ability of Muscovy, Pekin and their hybrids, hinny and mule ducks. Product. Animales. 2005;18:131–141. [Google Scholar]
  7. Bashir S., Qamar U., Khan F.H. Heterogeneous classifiers fusion for dynamic breast cancer diagnosis using weighted vote based ensemble. Qual. Quant. 2015;49:2061–2076. [Google Scholar]
  8. Benaddi H., Jouhari M., Elharrouss O. A lightweight hybrid approach for intrusion detection systems using a chi-square feature selection approach in IoT. Internet Things. 2025;32 [Google Scholar]
  9. Boateng E.Y., Otoo J., Abaye D.A. Basic tenets of classification algorithms K-nearest-neighbor, support vector machine, random forest and neural network: a review. J. Data Anal. Inform. Process. 2020;8:341–357. [Google Scholar]
  10. Bourouba R.A., Ghanem K.., Layeb A. Pages 1–5 in 2025 International Conference on Intelligent Computer Systems, Data Science and Applications (IC2SDA) 2025. A geometric extension of KNN classifier. [Google Scholar]
  11. Breiman L. Bagging predictors. Mach. Learn. 1996;24:123–140. [Google Scholar]
  12. Breiman L. Random forests. Mach. Learn. 2001;45:5–32. [Google Scholar]
  13. Chen J.T., He P..G., Jiang J.S., Yang Y.F., Wang S.Y., Pan C.H., Zeng L., He Y.F., Chen Z.H., Lin H.J., Pan J.M. In vivo prediction of abdominal fat and breast muscle in broiler chicken using live body measurements based on machine learning. Poult. Sci. 2023;102:102239. doi: 10.1016/j.psj.2022.102239. [DOI] [PMC free article] [PubMed] [Google Scholar]
  14. Chen Y.W., Chen K.Y. Vol. 11. China Poultry Industry News, Yangzhou, China; 2007. (Develop China’s Duck Fatty Liver Industry). [Google Scholar]
  15. Coşkun G., Şahin Ö., Altay Y., Aytekin İ. Final fattening live weight prediction in Anatolian merinos lambs from some body characteristics at the initial of fattening by using some data mining algorithms. Black Sea J. Agricult. 2023;6:47–53. [Google Scholar]
  16. Csóka Á., Simon S.E., Farkas T.P., Szász S., Sütő Z., Petneházy Ö., Kovács G., Repa I., Donkó T. In vivo estimation of chicken breast and thigh muscle weights using multi-atlas-based elastic registration on computed tomography images. Br. Poult. 2025;66(5):1–7. doi: 10.1080/00071668.2025.2472903. [DOI] [PubMed] [Google Scholar]
  17. Cui S., Wang D., Wang Y., Yu P.-W., Jin Y. An improved support vector machine-based diabetic readmission prediction. Comput. Methods Programs Biomed. 2018;166:123–135. doi: 10.1016/j.cmpb.2018.10.012. [DOI] [PubMed] [Google Scholar]
  18. Deng Y., Qu H., Leng A., Tang X., Zhai S. Methods and challenges in computer vision-based livestock anomaly detection, a systematic review. Biosyst. Eng. 2025;253 [Google Scholar]
  19. Dasarathy, B. V, and B. V Sheela. 1979. A composite classifier system design: Concepts and methodology. Proceedings of the IEEE 67:708–713.
  20. Feng Y., Wang X., Zhang J. A heterogeneous ensemble learning method for neuroblastoma survival prediction. IEEe J. Biomed. Health Inform. 2022;26:1472–1483. doi: 10.1109/JBHI.2021.3073056. [DOI] [PubMed] [Google Scholar]
  21. Flament A., Delleur V., Poulipoulis A., Marlier D. Corticosterone, cortisol, triglycerides, aspartate aminotransferase and uric acid plasma concentrations during foie gras production in male mule ducks (Anas platyrhynchos × Cairina moschata) Br. Poult. Sci. 2012;53:408–413. doi: 10.1080/00071668.2012.711468. [DOI] [PubMed] [Google Scholar]
  22. García-Vázquez F.A. Artificial intelligence and porcine breeding. Anim. Reprod. Sci. 2024;269 doi: 10.1016/j.anireprosci.2024.107538. [DOI] [PubMed] [Google Scholar]
  23. Guémené D., Guy G. The past, present and future of force-feeding and “foie gras” production. Worlds. Poult. Sci. J. 2004;60:210–222. [Google Scholar]
  24. Hamadani A., Ganai N.A., Mudasir S., Shanaz S., Alam S., Hussain I. Comparison of artificial intelligence algorithms and their ranking for the prediction of genetic merit in sheep. Sci. Rep. 2022;12 doi: 10.1038/s41598-022-23499-w. [DOI] [PMC free article] [PubMed] [Google Scholar]
  25. Hsu K.-W., Srivastava J. Springer; Berlin, Heidelberg: 2009. Diversity in Combinations of Heterogeneous Classifiers.Pages 923–932 in Advances in Knowledge Discovery and Data Mining. [Google Scholar]
  26. Huang Z., Yang C., Zhou X., Huang T. A hybrid feature selection method based on binary state transition algorithm and ReliefF. IEEe J. Biomed. Health Inform. 2019;23:1888–1898. doi: 10.1109/JBHI.2018.2872811. [DOI] [PubMed] [Google Scholar]
  27. Huma Z.E., Iqbal F. Predicting the body weight of Balochi sheep using a machine learning approach. Turk. J. Vet. Anim. Sci. 2019;43:500–506. [Google Scholar]
  28. Jia J.P., He X.Q., Jin Y.J. China Renmin University Press; Beijing, China: 2007. Statistics. [Google Scholar]
  29. Mavroforakis M.E., Theodoridis S. 2005 13th European Signal Processing Conference. 2005. Support Vector machine (SVM) classification through geometry; pp. 1–4. [Google Scholar]
  30. Mota L.F.M., Pegolo S., Baba T., Peñagaricano F., Morota G., Bittante G., Cecchinato A. Evaluating the performance of machine learning methods and variable selection methods for predicting difficult-to-measure traits in Holstein dairy cattle using milk infrared spectral data. J. Dairy. Sci. 2021;104:8107–8121. doi: 10.3168/jds.2020-19861. [DOI] [PubMed] [Google Scholar]
  31. National Standardization Administration . State Administration for Market Regulation; Beijing, China: 2024. Technical Specification for Determination of Production Performance of Meat Ducks (GB/T 29389-2024) [Google Scholar]
  32. Odhiambo Omuya E., Okeyo G.O., Kimwele M.W. Feature selection for classification using principal component analysis and information gain. Expert. Syst. Appl. 2021;174 [Google Scholar]
  33. Pan Q., Tang W., Yao S. The application of LightGBM in Microsoft malware detection. J. Phys. 2020;1684 [Google Scholar]
  34. Patil T.R., Sherekar S.S. Performance analysis of Naive Bayes and J48 classification algorithm for data classification. Int. J. Comput. Sci. Applic. 2013;6:256–261. [Google Scholar]
  35. Pilo B., George J.C. Diurnal and seasonal variation in liver glycogen and fat in relation to metabolic status of liver and M. pectoralis in the migratory starling, Sturnus roseus, wintering in India. Comp. Biochem. Physiol. a Physiol. 1983;74:601–604. doi: 10.1016/0300-9629(83)90554-6. [DOI] [PubMed] [Google Scholar]
  36. Pirompud P., Sivapirunthep P., Punyapornwithaya V., Chaosap C. Predictive modeling of bruising in broiler chickens using machine learning algorithms. Poult. Sci. 2025;104:105756. doi: 10.1016/j.psj.2025.105756. [DOI] [PMC free article] [PubMed] [Google Scholar]
  37. Schaeffer M.S., Levitt E.E. Concerning Kendall’s tau, a nonparametric correlation coefficient. Psychol. Bull. 1956;53:338. doi: 10.1037/h0045013. [DOI] [PubMed] [Google Scholar]
  38. Schapire R.E. The boosting approach to machine learning: an overview. Nonlinear Estim. Classific. 2003;171:149–171. [Google Scholar]
  39. Sesmero M.P., Ledezma A..I., Sanchis A. Generating ensembles of heterogeneous classifiers using stacked generalization. Wiley. Interdiscip. Rev. Data Min. Knowl. Discov. 2015;5:21–34. [Google Scholar]
  40. Sigletos G., Paliouras G., Spyropoulos C.D., Hatzopoulos M., Cohen W. Combining information extraction systems using voting and stacked generalization. J. Machine Learn. Res. 2005;6:1751–1782. [Google Scholar]
  41. Tao Y., Chen Y.J., Xue L., Xie C., Jiang B., Zhang Y. An ensemble model with clustering assumption for warfarin dose prediction in Chinese patients. IEEe J. Biomed. Health Inform. 2019;23:2642–2654. doi: 10.1109/JBHI.2019.2891164. [DOI] [PubMed] [Google Scholar]
  42. Thongkam J., Xu G., Zhang Y. 2008 IEEE international joint conference on neural networks (IEEE world congress on computational intelligence) IEEE; 2008. AdaBoost algorithm with random forests for predicting breast cancer survivability; pp. 3062–3069. [Google Scholar]
  43. Tırınk C., Piwczyński D., Kolenda M., Önder H. Estimation of body weight based on biometric measurements by using random forest regression, support vector regression and CART algorithms. Animals. 2023;13:798. doi: 10.3390/ani13050798. [DOI] [PMC free article] [PubMed] [Google Scholar]
  44. Trehiou S., Atallah E., Alquier-Bacquie V., Lasserre F., Arroyo J., Molette C., Remignon H. Development of hepatic steatosis in normal and veinous livers of overfed female mule ducks. Animal. 2025;19 doi: 10.1016/j.animal.2025.101502. [DOI] [PubMed] [Google Scholar]
  45. Vali Y., Gumpenberger M., Konicek C., Bagheri S. Computed tomography of the spleen in chickens. Frontiers in Veterinary Science. 2023;Volume:10–2023. doi: 10.3389/fvets.2023.1153582. [DOI] [PMC free article] [PubMed] [Google Scholar]
  46. Velusamy D., Ramasamy K. Ensemble of heterogeneous classifiers for diagnosis and prediction of coronary artery disease with reduced feature subset. Comput. Methods Programs Biomed. 2021;198 doi: 10.1016/j.cmpb.2020.105770. [DOI] [PubMed] [Google Scholar]
  47. Wang C., Chen X., Du L., Zhan Q., Yang T., Fang Z. Comparison of machine learning algorithms for the identification of acute exacerbations in chronic obstructive pulmonary disease. Comput. Methods Programs Biomed. 2020;188 doi: 10.1016/j.cmpb.2019.105267. [DOI] [PubMed] [Google Scholar]
  48. Wu X.X., Xiao C.B. Chongqing University Press; Chongqing, China: 2023. Animal Anatomy. [Google Scholar]
  49. Xu L., Duanmu Y., Blake G.M., Zhang C., Zhang Y., Brown K., Wang X., Wang P., Zhou X., Zhang M., Wang C., Guo Z., Guglielmi G., Cheng X. Validation of goose liver fat measurement by QCT and CSE-MRI with biochemical extraction and pathology as reference. European Radiology. 2018;28(5):2003–2012. doi: 10.1007/s00330-017-5189-x. [DOI] [PubMed] [Google Scholar]
  50. Yang P., Hwa Yang Y., Zhou B.B., Zomaya A.Y. A review of ensemble methods in bioinformatics. Curr. Bioinform. 2010;5:296–308. [Google Scholar]
  51. Yu K., Xie X. Predicting hospital readmission: a joint ensemble-learning model. IEEe J. Biomed. Health Inform. 2019;24:447–456. doi: 10.1109/JBHI.2019.2938995. [DOI] [PubMed] [Google Scholar]
  52. Zefrehi H.G., Altınçay H. Imbalance learning using heterogeneous ensembles. Expert. Syst. Appl. 2020;142 [Google Scholar]
  53. Zhang F.J., Liu Q.H., Qu H.L. Technical points for high-quality duck foie gras production. Proceedings of the 6th (2015) China Waterfowl Development Conference. Jilin Zhengfang Agriculture and Animal Husbandry Co., Ltd; China; National Waterfowl Industry Technology System Tonghua Experimental Station; 2015. pp. 257–258. [Google Scholar]
  54. Zhao X., Zhao Y., Zhang Y., Fan Q., Ke H., Chen X., Jin L., Tang H., Jiang Y., Ma J. Unraveling pathogenesis, biomarkers and potential therapeutic agents for endometriosis associated with disulfidptosis based on bioinformatics analysis, machine learning and experiment validation. J. Biol. Eng. 2024;18:42. doi: 10.1186/s13036-024-00437-0. [DOI] [PMC free article] [PubMed] [Google Scholar]
  55. Zhou X., Guan R., Cai H., Wang P., Yang Y., Wang X., Li X., Song H. Machine learning based personalized promotion strategy of piglets weaned per sow per year in large-scale pig farms. Porcine Health Manage. 2022;8:37. doi: 10.1186/s40813-022-00280-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
  56. Zhou Z.H. Tsinghua University Press; Beijing, China: 2016. Machine Learning. [Google Scholar]
  57. Zhou W.A. Beijing University of Posts and Telecommunications Press; Beijing, China: 2021. Machine Learning. [Google Scholar]
  58. Zhu W., Ma L., Shi Z., Qiao Y., Li Q., Pan B., Feng Z., Yang X., Cai J., Bai J., Sun L. Early-stage fertilised egg viability detection based on machine vision. Br. Poult. Sci. 66(5), 2025:1–12. doi: 10.1080/00071668.2025.2470275. [DOI] [PubMed] [Google Scholar]
  59. Zhu Z.P., Wang J., Gong D.Q., Duan X.J. Correlation analysis between body weight, body size and foie gras production performance of black-feathered muscovy ducks. Jiangsu Agricult. Sci. 2013;41:157–158. [Google Scholar]
  60. Zou K.H., Tuncali K.., Silverman S.G. Correlation and simple linear regression. Radiology. 2003;227:617–628. doi: 10.1148/radiol.2273011499. [DOI] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

mmc1.zip (293.8KB, zip)

Articles from Poultry Science are provided here courtesy of Elsevier

RESOURCES