Highlights
Analysis of humeral measurements from a contemporary French sample.
Comparison of predictive accuracy of seven statistical models.
Classification accuracy was greater than 90% for all models without selection methods.
Methods of variables selection was improved the accuracy of the models.
Relevance of new variables.
Penalized logistic regression (PLR), random forest (RF) and linear discriminant analysis (LDA) were the best accuracy models.
Keywords: Forensic anthropology, Sexual dimorphism, Sex estimation, Humerus bone, Statistical models, Machine learning
Abstract
Sex estimation is an important part of skeletal analysis and forensic identification. Traditionally pelvic traits are utilized for accurate sex estimation. However, the long bones, especially humerus, have been proved to be as effective for determine the sex of the individual.
The aim of this study was to compare the predictive accuracy of seven statistical modelling techniques including classical statistical methods and machine learning algorithms, to assess the sexual dimorphism of humerus on a French sample based on a metric analysis of 26 measurements. A total of 98 humeral bones (divided in two samples) were measured. Seven statistical models were compared: Linear Discriminant Analysis (LDA), Regularized Discriminant Analysis (RDA), Penalized Logistic Regression (PLR), Flexible Discriminant Analysis (FDA), Support Vector Machine (SVM), and Artificial Neural Network (ANN) and Random Forest (RF).
With cross validation, classification accuracy was greater than 90% (ranges between 92% and 98%) for all models without variable selection methods. The simplification of the models has improved the accuracy between 98% and 100% and also a reduction of the number of variables to 6 or less. Penalized logistic regression (PLR), Random Forest (RF) and Linear discriminant analysis (LDA) were the best accuracy models.
The measurements made at the proximal part of the humerus (WTT, CSD), at distal part (BEW, WT, MAW, THT) and of the entire bone (PLCT) stand out among the various models.
The present study suggests that the humerus is an interesting alternative for sex estimation and that non-classical statistical models can provide a new approach.
Introduction
Forensic anthropology is defined as a field of physical or biological anthropology that applies the anthropological methods for the examination and analysis of unknown human skeletal remains to establish the biological profile (sex, age, biogeographical origin, stature) and determine the causes and circumstances of death in a forensic context [1–4].
Sex estimation is the one of most important procedure for developing the biological profiles because estimation of age and stature are based on sex. Sex estimation can increase the possibility of human identification by 50% [e.g., 4–4]. The pelvis is considered the most accurate bone for sex estimation with correct classifications of at least 95% [e.g., 8–10]. The sexual dimorphism of the pelvis results from different reproductive roles between males and females [1]. However, post mortem damage and taphonomic changes may prevent the collection and subsequent analysis of this anatomical region. Traditionally, the skull was considered the second-best indicator for sex assessment with an accuracy of about 75 to 90% in forensic cases [e.g., 11]. It has now been established that the use of long bones, especially the upper limb bones, provides better results [e.g. 4, 12–15].
Forensic literature has reported the use of humerus in sex estimation [e.g., 16–26] with classification percentages frequently in excess of 80%, or even 90% in more recent studies. A variation in the degree of sexual dimorphism is reported among different population [15, 19] related to the interaction of many factors such as body size, genetics, hormonal status and by the time span or biogeographical origin [19, 21].
In fact, the humerus is the largest bone in the upper limb. Because of its structure and size, it is considered one of the strongest long bones in the skeleton and is one of the long bones that have been shown to remain in better condition after death. Moreover, even in a fragmented state, it is possible to recover information from it [13, 27, 28].
Different anthropological methods have been employed about humerus dimorphism such qualitative (non-metric) morphological [e.g. 29], quantitative (metric (single, combined, index measurements) [4, 7, 19–25, 31], or semi-quantitative (scoring). More recently morpho-geometric (a modelling system that allows the combination of both methods i.e. combining shape and size in 2D or even 3D) have been proposed when traditional methods are not feasible due to excessive fragmentation of the humerus [32, 33]. Most of the humerus ‘s studies using quantitative metric [4, 7, 19–25, 30, 31] or radiographic [26, 34] methods have shown good results with correct sex assessment accuracies ranging from 68–100%. However, there are differences concerning the most discriminating measures. For some authors, measurements taken on the proximal end of the humerus are more relevant [20, 22, 25, 26]. For others, it would be the total length [17, 23, 31]. Finally, some authors have established a greater dimorphism based on measurements of the distal end [24, 32]. The combination of several measurements would show higher sensitivity results [4, 22, 23, 25, 26, 34]. Indeed, several studies show that multiple measurements and multivariate techniques offer greater validity to the biological profile assessment.
To our knowledge, the forensic literature provides rare studies using classical statistical models (except for linear or flexible discriminant analysis frequently used) based on humerus measurements [4, 12, 21–23, 26, 34]. As far as artificial intelligence models are concerned, more and more studies have been using them in various forensic fields, particularly since 2015 [35]. In the forensic anthropology field of research, AI applications have been proposed by several authors for sex estimation [14, 35–37]. Only one study [37] was found concerning the use of nonlinear methods of discriminant analysis or machine learning classifiers on humerus. However, good model accuracy was reported when humeral measurements were combined with other long bone measurements (humerus, radius, femur, tibia). When humeral measurements were used in isolation, the percentage of correct classification did not exceed 90%. Moreover, only five measurements were made on the humeral bone (maximum length, epicondylar breadth, head diameter, diaphyseal mediolateral diameter, diaphyseal anteroposterior diameter at mid-shaft) and five statistical models were tested. Finally, sex was not always known with certainty, since it could be estimated.
Therefore, the objective of the present preliminary study was to compare the predictive accuracy of seven statistical modeling techniques, including machine learning algorithms, for assessing humeral sexual dimorphism only, using a larger number of variables, in a French sample.
Materials and methods
Bone samples
Bones used in this experiment were acquired from French individuals (European origin) that have donated their bodies to the science between 2004 and 2019. A total of 98 left humerus bones were analyzed. A first sample, called training sample composed by 60 humerus (30 males and 30 females) (S1) were utilized to fit the models of sex estimation. A second sample, test sample, comprising 38 humerus (18 males and 20 females) (S2) were utilized in order to verify the performance of the sex estimation models. Bones were randomly distributed between the reference and test samples. In the case of body donations to science, the sex was known with certainty. A hot water maceration was employed to remove the soft tissues and obtain clean specimens. Bones with macroscopic pathological or degenerative disorders were excluded.
Humerus measurements
The measurements were made on left humerus because it is assumed that handedness have no impact on bone size. Therefore, in any one individual, the measurements of a variable may theoretically be substituted by those of the other side depending on availability. In the literature, some studies found no significant difference in measurement according to the side [22, 24]. A total of 26 measurements were achieved. Measurements definitions were summarized in Table 1. Twenty-one of these measurements have already been described in the forensic literature. Five out of 26 were created for this study: Width between the Two Tubercles (WTT), Width of the Greater Tubercle (WGT), Width of the Lesser Tubercle (WLT), Width of Olecranon Fossa (WOF) and Thickness of Hollow of the Trochlea (THT) (Fig. 1). Five indices were also calculated from the measurements taken (Table 1).
Table 1.
Measurements selected for this study
| Variable | Brief definition | Measurement |
|---|---|---|
| Maximum length (ML) | Distance between the most proximal point of the humeral head and the most distal point of the trochlea, in posterior view | Osteometric board |
| Physiological length of medial epicondyle (PLME) | Distance between the most proximal point of the humeral head and the most distal point of the medial epicondyle, in anterior view | Osteometric board |
| Physiological length of lateral epicondyle (PLLE) | Distance between the most proximal point of the humeral head and the most distal point of the lateral epicondyle, in anterior view | Osteometric board |
| Physiological length at the hollow of the trochlea (PLHT) | Distance between the proximal articular surface of the humeral head and distal articular surface of the hollow of the trochlea, in anterior view | Small Spreading Caliper |
| Physiological length - hollow between capitulum and trochlea (PLCT) | Distance between the articular proximal surface of the humeral head and the distal articular surface of the hollow located between the capitulum and the trochlea, in anterior view | Small Spreading Caliper |
| Transversal diameter of the head (TDH) | Distance between the superior and inferior point of the head parallel to the head, in anterior view | Digital caliper |
| Vertical diameter of the head (VDH) | Distance between the anterior and posterior point of the head perpendicular to the long axis of the humerus, in medial view | Digital caliper |
| Diameter of the head with greater tubercle (DHGT) | Distance between the medial point of the articular surface of the head and the lateral point of the major tubercle, perpendicular to the long axis of the humerus, in anterior view | Digital caliper |
| Chirurgical neck circumference (CNC) | Perimeter of the neck at the junction between the head and the body of the humerus (constriction of the humerus located inferior to the greater and lesser tubercles) parallel to the long axis of the humerus, in anterior view | Measuring tape |
| Anatomical neck circumference (ANC) | Perimeter taken at the level of the zone that separates the head and the two tubercles, parallel to the head, in anterior view | Measuring tape |
| Circumference at sub deltoid level (CSD) | Perimeter taken at level of deltoid perpendicular to the long axis of the humerus | Measuring tape |
| Width between the two tubercles (WTT) | Distance from the most anterior point of the lesser tubercle to the most posterior point of the greater tubercle, perpendicular to the long axis of the diaphysis, in lateral view | Digital caliper |
| Width of the greater tubercle (WGT) | Distance between the most anterior and the most posterior point of the greater tubercle, in lateral view | Digital caliper |
| Width of the lesser tubercle (WLT) |
Distance between the lateral and medial edge of the lesser tubercle, perpendicular to the long axis of the diaphysis, in anterior view |
Digital caliper |
| Maximum diameter at mid height (MDMH) | Distance between the lateral and medial edges of the diaphysis perpendicular to the long axis of the humerus, at half of maximum length, in anterior view | Digital caliper |
| Minimum diameter at mid height (MinDMH) | Distance between the anterior and posterior edges of the diaphysis perpendicular to the long axis of the humerus, at half of maximum length, in lateral view | Digital caliper |
| Circumference at mdi height (CMH) | Perimeter taken at half of maximum length perpendicular to the long axis of the humerus | Measuring tape |
| Maximum diameter of the diaphysis (MDD) | Maximum diameter of the diaphysis approximately at the level of the lower third of the diaphysis, perpendicular to the long axis, in posterior view | Digital caliper |
| Minimum diameter of the diaphysis (MinDD) | Maximum diameter of the diaphysis approximately at the level of the lower third of the diaphysis, perpendicular to the long axis of the humerus, in posterior view | Digital caliper |
| Bi epicondylar width (BEW) | Distance between the most medial point of the medial epicondyle and the most lateral point of the lateral epicondyle, in posterior view | Digital caliper |
| Articular width (AW) | Distance from the lateral border of the capitulum to the medial border of the trochlea, in anterior view | Digital caliper |
| Maximum articular width (MAW) | Distance between the medial border of the trochlea and lateral epicondyle, in anterior vue | Digital caliper |
| Width of olecranon fossa (WOF) | Distance between the medial an lateral border of the fossa, in posterior view | Digital caliper |
| Width of the throclea (WT) | Distance between the medial an lateral border of the throclea, in anterior view | Digital caliper |
| Thickness of hollow of the trochlea (THT) | Distance between the most anterior and the most posterior point of the hollow of the trochlea, in lateral view | Digital caliper |
| With of capitulum (WC) | Distance between the medial and lateral border of the capitulum, in anterior view | Digital caliper |
| Robustness index at sub deltoid level (RISD) | Ratio between the circumference at sub deltoid level and maximum length | - |
| Robustness index at mid height (RIMH) | Ratio between the circumference at mid height and maximum length | - |
| Related index 1 (RI1) | Ratio between the maximum diameter at mid height and maximum length | - |
| Related index 2 (RI2) | Ratio between the minimum diameter at mid height and maximum length | - |
| Index of the diaphysis (ID) | Ration between the minimum diameter at mid height and the maximum diameter at mid height * 100 | - |
Fig. 1.
New variables in our study
Statistical analysis
Statistical analysis was performed using the RStudio software (version 4.2.0). Student T Test and Mann-Whitney test were used to compare each variable between males and females for S1 and S2 respectively (statistical significance p < 0.05).
Seven statistical prediction models were used in the present study and comprise models of discriminant analyses, logistic regression, and machine learning algorithms. The discriminant analysis included: Linear Discriminant Analysis (LDA), Flexible Discriminant Analysis (FDA) and Regularized Discriminant Analysis (RDA); the logistic regression was the Penalized Logistic Regression; and the machine learning include Artificial Neural Networks (ANN), Random Forest (RF) and Support Vector Machine (SVM).
Discriminant analyses models are used to explain and predict the probability of belonging to a group or category from measured data [38]. In our study, this statistical approach allows not only to predict sex, but also to estimate an individual’s probability of being male or female [39]. During the process of discriminant analysis, a value called centroid is calculated for each sex which is defined as the average value of all variables involved in the analysis for a sex group in a multivariate model. The position of the centroid is crucial in the model: for example, if the value of an individual is close to the position of the centroid, the probability that the individual belongs to this sex group is high. As a result of these calculations, a formula called discriminant function is obtained. When the values of an individual are placed in this function, the sex can be estimated depending on whether the result is larger or smaller than the cut-off point [40]. LDA is the parametric model of discriminant analysis [41]. It has been the most used model for sex estimation in forensic anthropology [e.g. 12, 19, 21, 42] and is almost the only model used in the literature to date for sex estimation from humeral bone measurements. This method recognizes a linear combination of predictor variables with separates mutually independent variables in an optimized way, and then creates a discriminant function that characterizes the differences between groups and classifies the new individuals whose group membership is not specified [43, 44]. It is a parametric classifier that models the probability of occurrence of one of two classes of a dichotomous dependent variable [44, 45]. Discriminant analysis is a multivariate parametric technique and, therefore, it is subject to several assumptions: the size of the smallest group must be greater than the number of sexual traits (independent variables), the variables must follow a multivariate normal distribution, the variance/covariance matrices of the variables must be homogeneous between the groups, while multi-collinearity between the variables must be excluded [39]. However discriminant analysis tolerates some weaknesses among the requirements, as usually it does not affect the performance of classification.
However, in the literature, these criteria were not always respected and do not affect the performance of the model. RDA and FDA are non-parametric versions of discriminant analysis [43]. FDA is a classification model based on a mixture of linear regression models, which uses optimal scoring to transform the response variable so that the data are in a better form for linear separation, and multiple adaptive regression splines to generate the discriminant surface [46, 47]. They result in a more flexible classifier that may show higher overall correct classification rates than LDA [45, 48].
Logistic regression models show the relationship between a set of continuous or categorical predictors X = (X1,., Xp) and a dichotomous outcome variable Y that represents the sex of individuals [41]. This model was excluding of the present study because of the difficulties in meeting the assumptions (distributions of our variables (normality not respected) and small sample size made it difficult to obtain a stable and reliable estimate). PLR is also one of the statistical models used to predict binary outcomes. The specificity of this model is to apply a penalty to the size of the L2 norm of the coefficients, decreasing the coefficients of the least contributing variables towards zero [40].
Machine Learning (ML) is based on data mining, allowing pattern recognition to provide predictive analysis. ML is a sub-category of artificial intelligence. It consists in letting algorithms discover “patterns”, that is recurring motifs, in data sets. These data can be numbers, words, images, statistics. By detecting patterns in these data, the algorithms learn and improve their performance in the execution of a specific task [49].
Different representations or models therefore learn to discriminate groups in different ways, so that different assumptions and constraints lead to different accuracy between algorithms depending on the problem [50]. ML models are not required to respect particular statistical assumptions, such as normality, since they are based on these data [50]. Indeed, they do not operate solely on group parameters such as means and covariance matrices. Instead, many algorithms must iteratively solve optimization problems in order to find the best set of neural weights (ANN) or tree structure (RF) that provides a decision boundary, often non-linear, with a maximum classification rate [50]. Several studies in the literature were particularly interesting as they demonstrated that traditional statistical classifiers can be easily outperformed by machine learning methods [50].
ANN is an artificial intelligence model corresponding to a system of interconnected neurons that mimics the functioning of the human brain [51–53]. It is one of the main tools used in ML and can be used for sex classification and prediction. Four basic elements compose this method: an input layer, a hidden layer, an output layer and synaptic connections between the above layers [40]. When ANN is used for sex classification, the leftmost layer of the network (input layer) consists of m input units that represent the traits X1, X2,…, Xm, while the output layer has only one node, which is related to the sex variable. In each neural network, the input values X1, X2,…, Xm are transformed into an output value via appropriate weights. A crush function (usually a logistic function or a hyperbolic tangent function) allows a non-linear transformation of the input. The layers are connected to each other, and the relationship between the nodes is represented by weights. The weights of the network are adjusted by a learning process. The main parameters for building an ANN are the number of hidden layers (size) and the decay parameter. The optimization of the ANN network involves the adjustment of the size and decay parameters.
RF is a non-parametric model that allows to explain a qualitative variable from one or more qualitative or quantitative explanatory variables. Random forests are a set of decision trees. RF predict a categorical dependent variable from measurements and observations on one or more predictor variables, through a series of rules, or nodes, similar to the branches of a tree [54].
SVM are non-probabilistic classifiers that are binary in their standard form [41]. SVM aim to find a linear hyperplane separation that maximizes the distance to the nearest of each class called margins. This model in not affected by the non-equality of group variances and for some authors, SVM is said to be the most robust and accurate method among all well-known algorithms [55].
The set of statistical prediction models was constructed in two steps. First, the reference sample (S1) was used to build the different models. All quantitative variables were included in the models in order to know the overall performance of each model in predicting the probability of belonging to the male or female sex class for each humerus. Some statistical models had methods for selecting the most relevant variables. The aim was twofold: to optimize the model by increasing its performance and to simplify it by reducing the number of variables to be entered into the model.
A selection variables method was also performed to improve the models and reduced the number of the variables. In LDA model a stepwise selection and a best subset selection were used. The stepwise selection is a simple and exhaustive selection method, which predicts the best variables using a bi-directional automated procedure that use forward selections, backward selections or both combined [56]. The best subset selection is exhaustive research for the best variables that identifies the best model containing a given number of predictors [56]. This method was also used in RDA model. For the other models (PLR, ANN, RF and SVM), selection variables are not required. However, for FDA and machine learnings it is possible to show the importance of the variables, in descending order, indicating the order of relevance of each variable to achieve the accuracy.
A cross-validation method (cross-validation, k = 10, repetition = 5) was used to prevent overfitting and artificial accuracy and allowed to obtain a minimum and maximum percentage of accuracies. A second sample (S2) was also used to assess the performance of the models. In the models with variable selection methods, only the variables selected in S1 were used. For the other methods all variables were used.
Results
Table 2 shows the statistical descriptive analysis of humerus measurements and the comparison between males and females. The measurements (average) obtained for males are higher than those obtained for females group. Statistical differences were observed in almost all variables with exception RI1 (p = 0.147, Mann Whitney test for independent series).
Table 2.
Statistical descriptive analysis of the entire sample (reference and test) and comparison between males and females (Mann Whitney test, independents series)
| Variables | Mean males (sd) | Mean females (sd) | p value (Wilcoxon test) |
|---|---|---|---|
| ML | 333.21 (11.34) | 301.58 (13.46) | < 0.001 |
| PLME | 323.38 (11.28) | 292.81 (13.06) | < 0.001 |
| PLLE | 322.52 (11.18) | 291.87 (12.70) | < 0.001 |
| PLHT | 327.91 (11.16) | 296.96 (13.51) | < 0.001 |
| PLCT | 328.24 (11.41) | 297.48 (12.76) | < 0.001 |
| TDH | 46.88 (2.23) | 41.15 (2.25) | < 0.001 |
| VDH | 50.95 (2.23) | 44.09 (2.23) | < 0.001 |
| DHGT | 53.95 (2.47) | 46.19 (2.97) | < 0.001 |
| CNC | 113.98 (5.54) | 96.00 (6.45) | < 0.001 |
| ANC | 164.08 (7.09) | 143.92 (5.91) | < 0.001 |
| CSD | 71.59 (4.51) | 62.15 (4.29) | < 0.001 |
| WTT | 49.69 (2.39) | 43.06 (2.56) | < 0.001 |
| WGT | 37.56 (2.64) | 32.58 (2.78) | < 0.001 |
| WLT | 20.15 (1.95) | 17.11 (1.82) | < 0.001 |
| MDMH | 23.60 (1.90) | 20.48 (1.76) | < 0.001 |
| MinDMG | 19.35 (1.37) | 16.13 (1.45) | < 0.001 |
| CMH | 72.46 (4.31) | 61.90 (4.59) | < 0.001 |
| MDD | 24.20 (1.48) | 21.27 (1.66) | < 0.001 |
| MinDD | 20.26 (1.50) | 16.87 (1.52) | < 0.001 |
| BEW | 65.14 (2.78) | 55.55 (2.94) | < 0.001 |
| AW | 47.55 (1.57) | 40.53 (2.51) | < 0.001 |
| MAW | 54.02 (2.29) | 45.94 (2.44) | < 0.001 |
| WOF | 25.51 (1.85) | 23.10 (1.67) | < 0.001 |
| WT | 26.62 (1.37) | 22.48 (1.38) | < 0.001 |
| THT | 18.42 (1.41) | 15.96 (1.20) | < 0.001 |
| WC | 18.48 (0.87) | 15.80 (1.07) | < 0.001 |
| RISD | 21.50 (1.38) | 20.47 (1.36) | 0.014 |
| RIMH | 21.96 (1.54) | 20.91 (1.45) | 0.026 |
| RI1 | 7.09 (0.60) | 6.80 (0.60) | 0.147 |
| RI2 | 5.81 (0.42) | 5.35 (0.45) | < 0.001 |
| ID | 82.11 (3.76) | 78.95 (5.48) | 0.024 |
Accuracy of different statistical models was presented in Table 3. The percentage of correctly classifications was greater than 90% using all variables. RDA was the best model among the discriminant analysis before variable selection. PLR displayed an accuracy average of 0.98. Accuracy of the different machine learnings ranges between 0.96 and 0.98. Variable selection methods were utilized to reduce the number of variables used in a final model to obtain a better performance of the different statistical models. In LDA model, accuracy increases from 0.92 to 0.98 when using stepwise method and to 0.99 with best subset selection. The number of variables used decreases to one and five (stepwise and best subset selection respectively). RDA and PLR models also increase the accuracy after selection methods. In the models without selection of variables (FDA, RF, ANN and SVM) we note that the percentage of correctly classifications is extremely high.
Table 3.
Accuracy (average) of the different statistical models under cross-validation
| Statistical model | Without selection methods (all variables included) (%) | With variables selected (%) |
Method of variables selection or optimization | Variables selected |
|---|---|---|---|---|
| LDA | 0.92 | 0.98 | Stepwise = (direction = “both”) | WTT |
| 0.99 | Best subset selection | PLCT, CSD, BEW, WT, RISD | ||
| RDA | 0.97 | 0.99 | Best subset selection | PLCT, CSD, BEW, WT, RISD |
| FDA | 0.96 | - | - | |
| PLR | 0.98 | 1 | Penalization of worst variables | PLLE, BEW, MAW, WT, WTT, THT |
| RF | 0.98 | - | - | |
| ANN | 0.96 | - | - | |
| SVM | 0.97 | - | - | |
Variables WTT (new variable), BEW and WT were presented in all almost selection methods (Table 3). Figures 2, 3, 4 and 5 to R show the importance of the variable’s contribution, in the models without variable selection methods (FDA, RF, ANN and SVM). In FDA model, only four variables are important to establish the accuracy of the model, WTT, DHGT, WT and PLLE (Fig. 2). If we make the prediction on the test sample with these four variables the accuracy goes up to one. In RF model (Fig. 3), there are five variables that are important more than 90% for the success of the model: WTT, DHGT, MAW, BEW and CNC. In ANN model, almost all variables are important to the performance of the model with exception of PLLE. RIMH is the only one that stands out (100% of importance) (Fig. 4). Concerning SVM model, there are 15 variables that are important than 90% (Fig. 5). Variables WTT, BEW, WT were presented in all seven statistical models. Variable DHGT is one of the best variables in the models without selection methods.
Fig. 2.
FDA’s plot
Fig. 3.
RF’s plot
Fig. 4.
ANN’s plot
Fig. 5.
SVM’s plot
Table 4 present the accuracy average of the different models on test sample (S2) with variables selected and using all variables for FDA, RF, ANN and SVM models. While some models predict 100% of the validation sample (LDA setpwise, PLR and RF) some models like LDA (best subset), RDA, and ANN drop off drastically and have poor predictions, close to 50%.
Table 4.
Accuracy of the different statistical models on test sample
| Statistical model | Accuracy test sample |
|---|---|
| LDA |
1 (stepwise) 0.53 (best subset selection) |
| RDA | 0.53 |
| FDA | 0.47 |
| PLR | 1 |
| RF | 1 |
| ANN | 0.65 |
| SVM | 0.84 |
Discussion
Sex estimation is an important part of analysis of skeletons and forensic identification process [8]. Traditionally pelvic traits are utilized for accurate sex estimation. However, the long bones, especially humerus bones, have been proved to be as effective for determine the sex of the individual [4, 12, 19]. The morphology of the humerus is considered to be a strong indicator of sex, particularly because women are thought to have narrower shoulders than men [17, 18]. It is also explained by the fact that this area is subject to greater biomechanical functioning and stresses and by the fact that the movements of the shoulder joint are greater than those of the elbow.
In the present study, 98 humeral bones (divided in two samples) were measured. As the scientific literature provides few studies concerning the use of non-linear discriminant analysis methods or machine learning classifiers in estimating sex from humeral bone [37] in comparison with classical discriminant analysis, the aim of this experiment was to compare the predictive accuracy of seven statistical modeling techniques, including machine learning algorithms, for assessing sexual dimorphisms of the humerus on a modern French sample. A total of 26 measurements were performed. Twenty-one variables have already been described in forensic literature. Five out of 26 were created for this study.
The descriptive statistics have showed that most of the measures were higher in men. The maximum values observed in females overlap all the minimum values observed in males as demonstrate in other studies [e.g., 19, 57]. This overlap zone is therefore unavoidable because there will always be short or slender men and tall or robust women. The comparison between males and females has showed significant differences except for the RI1 (Table 2) in agreement with the forensic literature [26, 58–61]. In fact, bone remodeling differs from males and females with more cortical bone development in male adolescence. In addition, maturation and bone growth in males is later when compared to females.
Reference sample
In the present study, seven statistical models were tested and the accuracy of which one was compared.
LDA model has produced an interesting accuracy of 92% using all variables. This accuracy was increased to 98% and 99% (depending on the variable selection method) (Table 3) with a drastic reduction on the number of variables. The stepwise method was allowed to select only one variable in LDA model (WTT). Most studies using discriminant analysis for sex determination already showed good results with accuracies of more than 80% [e.g. 21–23, 37, 62] and LDA is generally considered as more powerful when compared to non-parametric variants of a discriminant analysis [12]. However, they were also sometimes less effective [63].
Non-parametric models (RDA, FDA, PLR) that have the advantage of avoiding assumptions that are sometimes difficult to meet were also used. For the RDA model, the accuracy was high from the outset, rising from 97 to 99% after the variable selection method. The objective for this model was primarily to reduce the number of selected variables from 32 to 5. PLR was the model of non-parametric discriminant analysis with the highest accuracy and was already distinguished from other statistical models in the literature [36].
Machine learning had already proven its effectiveness in various forensic fields [35]. Thus, several studies have already indicated the potential of RF in anthropological research [e.g., 14, 37, 54]. The RF model has been described as a more efficient method than other commonly used techniques (discriminant analysis or logistic regression) [50], and it was this model that gave the highest percentage of correct classifications in the previous study on the humerus [37]. This algorithm allows to use the data to its full potential and to make objective and informed decisions. Classification trees have the advantage that they perform very well with missing data. Unlike discriminant analyses that must eliminate individuals missing even a single predictor variable, randomization forests can either remove the individual from a particular tree (or in some cases from a particular node) or use the median of the class estimate to replace the missing value; or use a proximity measure to extrapolate the missing data. Thus, not only do RFs perform better than discriminant analyses, but they also include a larger percentage of the original reference sample. Furthermore, it is interesting to note that some authors have shown that even when sample size and/or variance inequality negatively affected classification rates in more traditional methods such as logistic regression and discriminant function, classification trees such as random forest models became more accurate in comparison [64]. Finally, the RF was the simplest model as there was no need to select variables. ANN showed interesting results (96% accuracy) on the reference sample. ANNs had been shown to correctly predict sex in the literature [35, 36]. The main limitation of the application of ANN is the risk of over-fitting the model. The accuracy of SVM was 97%. SVM had also been shown to correctly predict sex in the literature [36, 37] with in some studies, such as ours, a better percentage of correct classification than ANNs [65]. The advantage of this algorithms is that you do not have to select the variables (however, you don’t know the variables used, only the order of importance).
In total, on the reference sample, the different models gave high accuracies, above 90%. For the models without variable selection (all variables included) the performance was above 95%. The variable selection allowed a drastic reduction of the number of variables.
Test sample
The performance of LDA with best subset selection, ANN, RDA and FDA models were strong in the training sample (S1) cross validated. However, in the test sample (S2) the models performed poorly with accuracy close to 50%. That can be explained by overfitting that occurred during cross-validation, where the models were highly tuned to the training data but failed to generalize to the test data. The predictive model gives very good predictions on Reference sample data (data it has already “seen” and adapted to), but will predict poorly on data it has not yet seen during its learning phase. This problem of overfitting, which led to average results in our study, contrasts with the more promising findings reported in some other studies [36]. Thus, in the test sample, the best performing models were LDA with stepwise, PLR and RF.
Consequently, in our study, the most efficient models were LDA (with stepwise selection), PLR and RF in both test samples (S1 and S2) (Tables 3 and 4). PLR and RF models gave the best results with an accuracy of 100% at the end of the variable selection method for PLR (98% with all the variables) and 98% for RF. On the test sample, an average performance was maintained at the end of the cross-validation, since whatever the 50 cross-validations, the accuracy obtained in the end for the two models was always 100%. These results agreed with a previous study on long bones using machine learning of where RF and PLR stands out with the highest accuracy and seems to be the best models [37]. Finally, and excepting the RF algorithm, these results support the notion that classical algorithms like LDA and PLR are still highly competitive and do not underperform when compared to machine learning based algorithms, particularly when the results need to be generalized to other samples.
Relevant variables
Selected variables differ according to the selection method that was used. Indeed, each variable selection method shows a group of best variables that are not necessarily the same from one method to another. This is not surprising, since each statistical model uses different variables to unlock the best that each model has to offer.
There would exist a tendency for the proximal segment to be more discriminating. Fasova [22] had hypothesized that proximal measurements would be more accurate because this area is subject to greater biomechanical and stress function while elbow movements would be more restricted than those of the shoulder. For the proximal epiphysis, the vertical diameter of the head would be the best parameter and offered a percentage of correct classification between 87% and 95.5% according to the authors [26, 34, 57–60, 66, 67], followed by the transverse diameter of the head [7, 20, 59], the humeral circumference [22], and the vertical diameter added to the width of the major tubercle [34]. In our study, it is indeed a variable measured on the proximal epiphysis, WTT, which appears to be very interesting, since it was selected in two of the statistical models (LDA and PLR), and even the only one selected for the LDA model with stepwise optimization, and made it possible to obtain a perfect average of 100% accuracy on the test sample at the end of the cross-validation. This variable had never been used before in the literature. On the other hand, the vertical diameter of the head and the transverse diameter of the head, the variables most cited in the literature, did not emerge from our algorithms. This could be explained by the fact that the populations from the studies cited were of different origin from ours and we know that there are important interpopulation differences in both bone morphology and bone size that do not allow extrapolation [68].
The distal end is also often cited in the literature. The most frequently cited variable is the epicondylar width, which is said to provide percentages of correct classification varying according to the authors from 68.8 to 91.62% [7, 22, 24, 26, 34, 57, 60, 69, 70]. In one study, the width of the olecranon fossa showed a high sexual dimorphism, with an accuracy of 94.0% [32]. Our results are thus compatible with those of the literature. Indeed, out of all our variables measured on the distal epiphysis, two variables stand out as being among the most relevant: BWE (in 3 models) and WT (in 3 models), offering a percentage of well classified cases, by being associated with other variables.
The maximum length of the humerus also provides good results in the literature, offering percentages of correct classification between 72.1 and 93.3% [7, 23, 26, 31, 34, 66]. Our results are therefore correlated. Indeed, PLCT (Physiological length of the proximal articular surface of the head - hollow between capitulum and trochlea) and PLLE (physiological length of the lateral epicondyle) are two of the 27 variables retained by our algorithms. LPCT is retained in two models and LPEL in one model, offering percentages of correct classification, in association with the other retained variables, between 97 and 100%.
Thus, in our study, the most discriminating variables could concern the proximal and distal epiphysis and the length of the bone, as observed by most authors.
Two variables targeting other anatomical areas, and an index stand out: PSD (CSD), RISD. CSD, rarely cited in the literature [31], is selected here in two of the algorithm models and offers a classification percentage of 99% in combination with the other selected variables. IRSD, which is the subdeltoidal robustness index, never cited in the literature, when co-selected with other variables, allowed us to obtain an excellent classification rate of 99%.
Several advantage and limitations must be mentioned concerning the sample. First, we used quantitative methods. Unlike the qualitative methods are based on the description of observational traits and are often very subjective depending on the experience of the examiner and subject to high intra- and inter-observer error, quantitative methods are based on measurements with precisely defined variables for greater simplicity, objectivity and reproducibility. Indeed, they allow the elimination of the subjectivity inherent to morphological evaluation and thus reduce inter-observer and intra-observer errors [18]. Secondly the bones were mainly from scientific donations, forensic autopsies, or anthropological expertise and all came from European individuals from the south of France. The sex was perfectly known for each individual and the population used was contemporary (recent). Indeed, variations on the bone according to the chronology called secular changes had been noted in the literature.
The small sample size was certainly perfectly characterized, but prevented generalization of the statistical models and thus constituted the main limitation of the study. In particular, models suffering from overfitting would have required a larger learning sample for better generalization.
A limit can be found in the fact that the sample was limited and came from elderly people not representative of the general population. Indeed, some studies had shown that age-related bone changes could lead to a misclassification of sex. Ageing is thought to influence skeletal morphology in different ways [29]. During the ageing process, bone resorption occurs on the cortical surface, while at the same time this process is compensated for by periosteal apposition and bone enlargement [71, 72]. Thus, young boy would tend to be classified as young girl and older women would tend to be classified as men. Thus, it is not possible to extrapolate our results to a general population and especially to the young subject. It therefore appears necessary to test these variables on a larger and more diversified sample (mixing subjects of different ages and origins) and renew the samples because the old populations are no longer a reflection of the subjects of our time.
In conclusion, in the present study, we compared the predictive accuracy of different statistical models for sex estimation generated from humeral bone measurements of a French sample.
The overall performance of the models is very similar, ranging from 92 to 98% or even 100% after correct classification on an independent sample and demonstrates the relevance of the humerus in sexual dimorphism. PLR and RF were the two models that stood out from the other statistical models with an average performance of 100%. RF was the simpler of the two since it did not need to use any method of variable selection. The most discriminating measures were for the proximal and distal ends of the humerus. Three variables appear to be the focus of this study: BWE and WT, which were already described as relevant in the literature, and WTT, a new variable never used before. The present study suggests that the humeral bone constitutes a valid alternative for sex estimation of skeletal remains with comparable classification accuracies to the pelvis or femur and that the non-classical statistical models may provide a novel approach to sex estimation from the humeral bone.
The relatively small sample size (n = 98) is a limitation of the study, but the sex of all bones was fully documented.
Funding
Open access funding provided by CHRU de Brest.
Data availability
Not applicable.
Code availability
Not applicable.
Declarations
Ethics approval
Compliance with ethical standards.
Consent to participate
Not applicable.
Consent to publish
Not applicable.
Clinical trial number
not applicable.
Conflict of interest
There is no conflict of interest of this work.
Footnotes
Publisher’s note
Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
References
- 1.Gray H, Lewis WH (1918) Anatomy of the human body. 10.5962/bhl.title.20311
- 2.Murail P, Bruzek J, Houët F, Cunha E (2005) DSP: a tool for probabilistic sex diagnosis using worldwide variability in hip-bone measurements. Bulletins et Mémoires de la Société D Anthropologie de Paris 17(3–4):167–176. 10.4000/bmsap.1157 [Google Scholar]
- 3.Phenice TW (1969) A newly developed visual method of sexing the os pubis. Am J Phys Anthropol 30(2):297–301. 10.1002/ajpa.1330300214 [DOI] [PubMed] [Google Scholar]
- 4.Spradley MK, Jantz RL (2011) Sex estimation in Forensic Anthropology: Skull Versus Postcranial Elements. J Forensic Sci 56(2):289–296. 10.1111/j.1556-4029.2010.01635.x [DOI] [PubMed] [Google Scholar]
- 5.Scheuer L (2002) Application of osteology to forensic medicine. Clin Anat 15(4):297–312. 10.1002/ca.10028 [DOI] [PubMed] [Google Scholar]
- 6.Durić M, Rakocević Z, Donić D (2005) The reliability of sex determination of skeletons from forensic context in the Balkans. Forensic Sci Int 147(2–3):159–164. 10.1016/j.forsciint.2004.09.111 [DOI] [PubMed] [Google Scholar]
- 7.Ferembach D, Schwidetzky I, Stloukal M (1979) Recommandations pour déterminer l’âge et le sexe sur le squelette. Bulletins et Mémoires de la Société D Anthropologie de Paris 6(1):7–45. 10.3406/bmsap.1979.1945 [Google Scholar]
- 8.Quatrehomme G (2015) Traité d’Anthropologie médico-légale. De Boeck Supérieur, Louvain-la-Neuve. (Belgique) [Google Scholar]
- 9.Krishan K, Chatterjee PM, Kanchan T, Kaur S, Baryah N, Singh RK (2016) A review of sex estimation techniques during exami- nation of skeletal remains in forensic anthropology casework. Forensic Sci Int 261:165.e1-165.e8 [DOI] [PubMed]
- 10.Franklin D, Cardini A, Flavel A, Marks MK (2014) Morphometric analysis of pelvic sexual dimorphism in a contemporary western Australian population. Int J Legal Med 28(5):861–872 [DOI] [PubMed] [Google Scholar]
- 11.Franklin L, Freedman N (2005) Sexual dimorphism and discriminant function sexing in indigenous South. Afr crania Homo 55:213–228 [DOI] [PubMed] [Google Scholar]
- 12.Krüger GC, L’Abbé EN, Stull KE (2016) Sex estimation from the long bones of modern South africans. Int J Legal Med 131(1):275–285. 10.1007/s00414-016-1488-z [DOI] [PubMed] [Google Scholar]
- 13.Bongiovanni R, LeGarde CB (2017) A Univariate Approach to Sex Estimation for the fragmentary Upper Limb. J Forensic Sci 63(2):356–360. 10.1111/1556-4029.13530 [DOI] [PubMed] [Google Scholar]
- 14.Nogueira L, Santos F, Castier F, Knecht S, Bernardi C, Alunni V (2023) Sex assessment using the radius bone in a French sample when applying various statistical models. Int J Legal Med 137(3):925–934. 10.1007/s00414-023-02981-8 [DOI] [PubMed] [Google Scholar]
- 15.Bidmos MA, Mazengenya P (2020) Accuracies of discriminant function equations for sex estimation using long bones of upper extremities. Int J Legal Med 135(3):1095–1102. 10.1007/s00414-020-02458-y [DOI] [PubMed] [Google Scholar]
- 16.Bass WM (2005) Human osteology: a Laboratory and Field Manual. 5e éd. Missouri Archaeological Society, Columbia, Mo [Google Scholar]
- 17.Rogers TL (1999) A visual method of determining the sex of skeletal remains using the distal humerus. J Forensic Sci 44(1):14411J. 10.1520/jfs14411j [PubMed] [Google Scholar]
- 18.Tanaka H, Lestrel PE, Uetake T, Kato S, Ohtsuki F (2000) Sex differences in proximal humeral outline shape: elliptical fourier functions. J Forensic Sci 45(2):14682J. 10.1520/jfs14682j [PubMed] [Google Scholar]
- 19.Charisi D, Eliopoulos C, Vanna V, Koilias CG, Manolis SK (2010) Sexual dimorphism of the Arm bones in a Modern Greek Population. J Forensic Sci 56(1):10–18. 10.1111/j.1556-4029.2010.01538.x [DOI] [PubMed] [Google Scholar]
- 20.Dittrick J, Suchey JM (1986) Sex determination of prehistoric central California skeletal remains using discriminant analysis of the femur and humerus. Am J Phys Anthropol 70(1):3–9. 10.1002/ajpa.1330700103 [DOI] [PubMed] [Google Scholar]
- 21.Mall G, Hubig M, Büttner A, Kuznik J, Penning R, Graw M (2001) Sex determination and estimation of stature from the long bones of the arm. Forensic Sci Int 117(1–2):23–30. 10.1016/s0379-0738(00)00445-x [DOI] [PubMed] [Google Scholar]
- 22.Fasova AV, Timonov PT (2017) Sex determination from the proximal and distal part of the humerus in a Bulgarian contemporary population. Anil Aggrawal’s Internet J Foren Med Toxicol [serial online] 18(1):10
- 23.Mokoena P, Billings BK, Gibbon V, Bidmos MA, Mazengenya P (2019) Development of discriminant functions to estimate sex in upper limb bones for mixed ancestry South africans. Sci Justice 59(6):660–666. 10.1016/j.scijus.2019.06.007 [DOI] [PubMed] [Google Scholar]
- 24.Attia MH, Aboulnoor BAE (2020) Tailored logistic regression models for sex estimation of unknown individuals using the published population data of the humeral epiphyses. Leg Med 45:101708. 10.1016/j.legalmed.2020.101708 [DOI] [PubMed] [Google Scholar]
- 25.Kranioti EF, Michalodimitrakis M (2009) Sexual dimorphism of the Humerus in contemporary Cretans—A Population-Specific Study and a review of the literature. J Forensic Sci 54(5):996–1000. 10.1111/j.1556-4029.2009.01103.x [DOI] [PubMed] [Google Scholar]
- 26.Kranioti EF, Nathena D, Michalodimitrakis M (2010) Sex estimation of the Cretan humerus: a digital radiometric study. Int J Legal Med 125(5):659–667. 10.1007/s00414-010-0470-4 [DOI] [PubMed] [Google Scholar]
- 27.Waldron T (1987) The relative survival of the human skeleton: implications for paleopathology. In: Boddington JRA, Garland AN (eds) Death, Decay Reconstr. Manchester University, Manchester, pp 55–64 [Google Scholar]
- 28.Galloway A, Snyder L, Willey P (1996) Human bone mineral densities and survival of bone elements. Dans CRC Press eBooks. 10.1201/9781439821923.ch19
- 29.Tallman SD, Blanton AI (2019) Distal Humerus Morphological Variation and Sex Estimation in Modern Thai individuals. J Forensic Sci 65(2):361–371. 10.1111/1556-4029.14218 [DOI] [PubMed] [Google Scholar]
- 30.Tise ML, Spradley MK, Anderson BE (2012) Postcranial sex estimation of individuals considered hispanic. J Forens Sci 58(s1). 10.1111/1556-4029.12006 [DOI] [PubMed]
- 31.Vaishnani HV, AR G, GV S (2019) A study on sexual dimorphism of the humerus in central gujarat. Int J Anat Res 7(23):6668–6673. 10.16965/ijar.2019.200 [Google Scholar]
- 32.Ammer S, Coelho JD, Cunha EM (2019) Outline shape analysis on the Trochlear Constriction and Olecranon Fossa of the Humerus: insights for sex estimation and a New Computational Tool. J Forensic Sci 64(6):1788–1795. 10.1111/1556-4029.14096 [DOI] [PubMed] [Google Scholar]
- 33.López-Lázaro S, Pérez-Fernández A, Alemán I, Viciano J (2020) Sex estimation of the humerus: a geometric morphometric analysis in an adult sample. Leg Med 47:101773. 10.1016/j.legalmed.2020.101773 [DOI] [PubMed] [Google Scholar]
- 34.Shehri FA, Soliman KE (2015) Determination of sex from radiographic measurements of the humerus by discriminant function analysis in Saudi population, Qassim region, KSA. Forensic Sci Int 253:138e1. 138.e6 [DOI] [PubMed] [Google Scholar]
- 35.Galante N, Cotroneo R, Furci D, Lodetti G, Casali MB (2022) Applications of artificial intelligence in forensic sciences: Current potential benefits, limitations and perspectives. Int J Legal Med 137(2):445–458. 10.1007/s00414-022-02928-5 [DOI] [PubMed] [Google Scholar]
- 36.Knecht S, Nogueira L, Servant M, Santos F, Alunni V, Bernardi C, Quatrehomme G (2021) Sex estimation from the greater sciatic notch: a comparison of classical statistical models and machine learning algorithms. Int J Legal Med 135(6):2603–2613. 10.1007/s00414-021-02700-1 [DOI] [PubMed] [Google Scholar]
- 37.Knecht S, Santos F, Ardagna Y, Alunni V, Adalian P, Nogueira L (2023) Sex estimation from long bones: a machine learning approach. Int J Legal Med 137(6):1887–1895. 10.1007/s00414-023-03072-4 [DOI] [PubMed] [Google Scholar]
- 38.Haeb-Umbach R, Ney H (1992) Linear discriminant analysis for improved large vocabulary continuous speech recognition. ICASSP-92: 1992 IEEE Int Conf Acoust Speech Signal Process. 10.1109/icassp.1992.225984 [Google Scholar]
- 39.Nikita E, Nikitas P (2020) Sex estimation: a comparison of techniques based on binary logistic, probit and cumulative probit regression, linear and quadratic discriminant analysis, neural networks, and naïve Bayes classification using ordinal variables. Int J Legal Med 134(3):1213–1225. 10.1007/s00414-019-02148-4 [DOI] [PubMed] [Google Scholar]
- 40.Etli Y, Asirdizer M, Hekimoglu Y, Keskin S, Yavuz A (2019) Sex estimation from sacrum and coccyx with discriminant analyses and neural networks in an equally distributed population by age and sex. Forensic Sci Int 303:109955. 10.1016/j.forsciint.2019.109955 [DOI] [PubMed] [Google Scholar]
- 41.Santos F, Guyomarc’h P, Bruzek J (2014) Statistical sex determination from craniometrics: comparison of linear discriminant analysis, logistic regression, and support vector machines. Forensic Sci Int 245:204e1. 204.e8 [DOI] [PubMed] [Google Scholar]
- 42.Stull KE, L’Abbé EN, Ousley SD (2017) Subadult sex estimation from diaphyseal dimensions. Am J Phys Anthropol 163(1):64–74. 10.1002/ajpa.23185 [DOI] [PubMed] [Google Scholar]
- 43.Coelho JD, Curate F (2019) CADOES: an interactive machine-learning approach for sex estimation with the pelvis. Forensic Sci Int 302:109873. 10.1016/j.forsciint.2019.109873 [DOI] [PubMed] [Google Scholar]
- 44.Curate F, Umbelino C, Perinha A, Nogueira C, Silva A, Cunha E (2017) Sex determination from the femur in Portuguese populations with classical and machine-learning classifiers. J Forensic Leg Med 52:75–81. 10.1016/j.jflm.2017.08.011 [DOI] [PubMed] [Google Scholar]
- 45.Hinić-Frlog S, Motani R (2010) Relationship between osteology and aquatic locomotion in birds: determining modes of locomotion in extinct Ornithurae. J Evol Biol 23(2):372–385. 10.1111/j.1420-9101.2009.01909.x [DOI] [PubMed] [Google Scholar]
- 46.Hallgren W, Santana F, Low-Choy S, Zhao Y, Mackey B (2019) Species distribution models can be highly sensitive to algorithm configuration. Ecol Model 408:108719. 10.1016/j.ecolmodel.2019.108719 [Google Scholar]
- 47.Hastie T, Tibshirani R, Friedman J (2009) The elements of statistical learning: data mining, inference, and prediction. Springer
- 48.Paliwal M, Kumar UA (2009) Neural networks and statistical techniques: a review of applications. Expert Syst Appl 36(1):2–17. 10.1016/j.eswa.2007.10.005 [Google Scholar]
- 49.Uddin S, Khan A, Hossain ME, Moni MA (2019) Comparing different supervised machine learning algorithms for disease prediction. BMC Med Inf Decis Mak 19(1). 10.1186/s12911-019-1004-8 [DOI] [PMC free article] [PubMed]
- 50.Navega D, Vicente R, Vieira DN, Ross AH, Cunha E (2014) Sex estimation from the tarsal bones in a Portuguese sample: a machine learning approach. Int J Legal Med 129(3):651–659. 10.1007/s00414-014-1070-5 [DOI] [PubMed] [Google Scholar]
- 51.Haykin SO (2007) Neural networks and learning machines (International edn)
- 52.Ripley BD (2007) Pattern recognition and neural networks. Cambridge University Press
- 53.Du Jardin P, Ponsaillé J, Alunni-Perret V, Quatrehomme G (2009) A comparison between neural network and other metric methods to determine sex from the upper femur in a modern French population. Forensic Sci Int 192(1–3):127.e1-127.e6. 10.1016/j.forsciint.2009.07.014 [DOI] [PubMed]
- 54.Hefner JT, Spradley MK, Anderson B (2014) Ancestry assessment using random forest modeling. J Forensic Sci 59(3):583–589. 10.1111/1556-4029.12402 [DOI] [PubMed] [Google Scholar]
- 55.Wu X, Kumar V, Quinlan JR, Ghosh J, Yang Q, Motoda H, McLachlan GJ, Ng A, Liu B, Yu PS, Zhou Z, Steinbach M, Hand DJ, Steinberg D (2007) Top 10 algorithms in data mining. Knowl Inf Syst 14(1):1–37. 10.1007/s10115-007-0114-2 [Google Scholar]
- 56.Maroco J, Silva D, Rodrigues A, Guerreiro M, Santana I, De Mendonça A (2011) Data mining methods in the prediction of dementia: a real-data comparison of the accuracy, sensitivity and specificity of linear discriminant analysis, logistic regression, neural networks, support vector machines, classification trees and random forests. BMC Res Notes 4(1). 10.1186/1756-0500-4-299 [DOI] [PMC free article] [PubMed]
- 57.Frutos LR (2005) Metric determination of sex from the humerus in a Guatemalan forensic sample. Forensic Sci Int 147(2–3):153–157. 10.1016/j.forsciint.2004.09.077 [DOI] [PubMed] [Google Scholar]
- 58.Sikka A, Jain A (2016) Sexual dimorphism in humerus: a morphometric study in the north Indian population. Int J Basic Appl Med Sci
- 59.Lee J, Kim Y, Lee U, Park D, Jeong Y, Lee NS, Han SY, Kim K, Han S (2014) Sex determination using upper limb bones in Korean populations. Anat Cell Biology 47(3):196. 10.5115/acb.2014.47.3.196 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 60.Devi S, Raichandani L, Kataria SK, Raichandani S (2013) A osteometric study of sex determination by epiphysial ends of humerus in western Rajasthan sample
- 61.Vaishnani HV, AR G, GV S (2019b) A study on sexual dimorphism of the humerus in central gujarat. Int J Anat Res 7(23):6668–6673. 10.16965/ijar.2019.200 [Google Scholar]
- 62.Ahmed SS, Siddiqui FB, Bayer SB (2018) Sex differentiation of Humerus: an osteometric study. J Clin Diagn Res. 10.7860/jcdr/2018/35768.12325 [Google Scholar]
- 63.Feldesman MR (2002) Classification trees as an alternative to linear discriminant analysis. Am J Phys Anthropol 119(3):257–275. 10.1002/ajpa.10102 [DOI] [PubMed] [Google Scholar]
- 64.Finch H, Schneider MK (2007) Classification Accuracy of Neural Networks vs. Discriminant Analysis, logistic regression, and classification and regression trees. Methodology 3(2):47–57. 10.1027/1614-2241.3.2.47 [Google Scholar]
- 65.Toneva D, Nikolova S, Agre G, Zlatareva D, Hadjidekov V, Lazarov N (2020) Machine learning approaches for sex estimation using cranial measurements. Int J Legal Med 135(3):951–966. 10.1007/s00414-020-02460-4 [DOI] [PubMed] [Google Scholar]
- 66.Moore MK, DiGangi EA, Ruíz F P N, Davila OJH, Medina CS (2016) Metric sex estimation from the postcranial skeleton for the Colombian population. Forensic Sci Int 262:286e1. 286.e8 [DOI] [PubMed] [Google Scholar]
- 67.Goncalves D (2014) Evaluation of the effect of secular changes in the reliability of osteometric methods for the sex estimation of Portuguese individuals. Cadernos do GEEvH 3(1)
- 68.Steyn M, İşcan M (1999) Osteometric variation in the humerus: sexual dimorphism in South africans. Forensic Sci Int 106(2):77–85. 10.1016/s0379-0738(99)00141-3 [DOI] [PubMed] [Google Scholar]
- 69.Spradley MK, Anderson BE, Tise ML (2014) Postcranial sex estimation criteria for Mexican hispanics. J Forens Sci 60(s1). 10.1111/1556-4029.12624 [DOI] [PubMed]
- 70.Boldsen JL, Milner GR, Boldsen SK (2015) Sex estimation from modern American humeri and femora, accounting for sample variance structure. Am J Phys Anthropol 158(4):745–750. 10.1002/ajpa.22812 [DOI] [PubMed] [Google Scholar]
- 71.Stini WA (1985) Growth rates and sexual dimorphism in evolutionary perspective. In: Gilbert RI, Mielke JH (eds) The analysis of prehistoric diets. Academic, Orlando, pp 191–226 [Google Scholar]
- 72.Curate F, Cunha E (2017) Femoral cortical bone in a Portuguese reference skeletal collection. Antropol Port 34:91–109. 10.14195/2182-7982_34_5 [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Data Availability Statement
Not applicable.
Not applicable.





