Skip to main content
NIHPA Author Manuscripts logoLink to NIHPA Author Manuscripts
. Author manuscript; available in PMC: 2024 Feb 1.
Published in final edited form as: Acad Radiol. 2022 Oct 3:S1076-6332(22)00495-0. doi: 10.1016/j.acra.2022.09.004

Multiparametric Quantitative Imaging Biomarkers for Phenotype Classification: A Framework for Development and Validation

Jana G Delfino 1, Gene A Pennello 2, Huiman X Barnhart 3, Andrew J Buckler 4, Xiaofeng Wang 5, Erich P Huang 6, Dave L Raunig 7, Alexander R Guimaraes 8, Timothy J Hall 9, Nandita NM deSouza 10, Nancy Obuchowski 11
PMCID: PMC9825632  NIHMSID: NIHMS1835352  PMID: 36202670

Abstract

This manuscript is the third in a five-part series related to statistical assessment methodology for technical performance of multi-parametric quantitative imaging biomarkers (mp-QIBs). We outline approaches and statistical methodologies for developing and evaluating a phenotype classification model from a set of multiparametric QIBs. We then describe validation studies of the classifier for precision, diagnostic accuracy, and interchangeability with a comparator classifier. We follow with an end-to-end real-world example of development and validation of a classifier for atherosclerotic plaque phenotypes. We consider diagnostic accuracy and interchangeability to be clinically meaningful claims for a phenotype classification model informed by mp-QIB inputs, aiming to provide tools to demonstrate agreement between imaging-derived characteristics and clinically established phenotypes. Understanding that we are working in an evolving field, we close our manuscript with an acknowledgement of existing challenges and a discussion of where additional work is needed. In particular, we discuss the challenges involved with technical performance and analytical validation of mp-QIBs. We intend for this manuscript to further advance the robust and promising science of multiparametric biomarker development.

Keywords: phenotype classification, multiparametric classification, multi-class classification, multi-parametric quantitative imaging biomarkers (mp-QIBs), QIBA

1. Introduction

Quantitative imaging biomarkers (QIBs) are objectively measured, ratio or interval scale characteristics derived from one or more in vivo images that indicate a normal biological process, a pathogenic process, or a response to a therapeutic intervention1. In a 2015 metrology series16 the Quantitative Imaging Biomarker Alliance (QIBA) reviewed statistical methods for characterizing the analytical / technical performance of QIBs. Multiple QIBs, when combined to form a multi-parametric QIB (mp-QIB), promise additional clinical utility over single quantitative imaging biomarkers for characterizing tissue, detecting disease, identifying phenotypes and longitudinal change, predicting outcomes, or other potential intended uses. Multiple QIBs can be selected from within single or multiple imaging modalities to represent different aspects of tissue morphology or biology. If selected carefully, each QIBs has the potential to contribute unique clinical information, and the value of the mp-QIB exceed that of its individual components.

In a new series of papers, QIBA identifies four use cases for mp-QIBs7. For each use case, a framework for the development and validation of the relevant mp-QIBs is described. In the first use case, a vector of p QIBs X = (X1, …, Xp) is treated as a p-dimensional multivariate descriptor of health8. This manuscript discusses the second use case: combining X1, …, Xp to differentiate between phenotypes. The third use case involves combining X1, …, Xp to forecast the risk of occurrence of an event of interest, such as death, disease progression, or recurrence9. Finally, the fourth use case describes the applicability of these frameworks to radiomic data10.

Phenotypes are observable characteristics of a subject, including those arising from response to the environment11. Statistically, a phenotype can be one of three types of categorical variables: (1) dichotomous, (2) nominal polychotomous--that is, more than two possible values that cannot be meaningfully ordered--and (3) ordinal polychotomous--that is, more than two possible values that can be meaningfully ordered but are not quantitative, i.e., not interval or ratio scale. In metrology and medicine, ordinal variables are called “semi-quantitative”. Examples of a dichotomous phenotype variable are presence or absence of a disease and benign or malignant classification of a lesion or nodule12. Examples of a nominal, polychotomous phenotype variable are stroke subtype13, skin lesion type14, and breast cancer subtype15. Examples of a semi-quantitative (ordinal) phenotype variable are liver fibrosis stage16, histopathological classification of cervical intraepithelial neoplasia17, and prostate Gleason grade18 (Table 1). While phenotypes can be identified from an individual feature, for example vascularity (through dynamic contrast enhanced studies), glycolytic metabolism (through standardized uptake values of glucose on 18F-FDG PET scans) or elasticity (through measurements of shear wave on ultrasound or MRI), the subject of this manuscript is combining multiple quantitative measurements into a defined phenotype.

Table 1 –

Examples of Phenotype Classification Tasks

Phenotype Classification Task Definition Examples
Dichotomous Two categories Malignant/benign classification
Presence/absence of disease
Nominal polychotomous Three or more categories, not ordered Atherosclerotic Plaque Type Classification (subclinical/stable/unstable)
Blood type (A, B, O, AB)
Breast lesion classification (benign, mass, calcification, architectural distortion, etc.)
Stroke subtype (large-artery atherosclerosis, cardioembolism, small-vessel occlusion, stroke of other determined etiology, stroke of undetermined etiology)
Lung Texture (normal, hyperlucency, fibrosis, ground glass, reticular, honeycomb, solid, focal, etc.)
Ordinal polychotomous Three or more categories, ordered Disease Severity or Staging, e.g. METAVIR liver fibrosis stage F0-F4, cancer staging (I-IV), prostate Gleason grade, Breast imaging reporting and database system (BI-RADS) score 1–6.
Head CT scan for intracranial hemorrhage, mass effect, midline shift, or cranial fracture
Cervical intraepithelial neoplasia (CIN) histology (normal, CIN 1, CIN 2, CIN 3, cancer)
Pap smear cytology (NILM, ASC-US, LSIL, HSIL, cancer)
Diagnostic tests with equivocal zones, e.g., hepatitis C test negative, equivocal, or positive.

Development of a phenotype classification model from imaging findings requires first demonstrating connection between an existing, stable phenotype and imaging findings (e.g. demonstrating biological plausibility). This discovery phase is followed by development and refinement of the model/classifier and assessment of the analytical/technical performance of the model/classifier. Evaluation of the clinical implications, sometimes termed clinical utility, of classification (e.g. prognosis prediction, treatment selection, etc.) is the fourth and final step. We focus this manuscript on model/classifier development and analytical/technical performance assessment.

We argue that diagnostic accuracy (defined here as agreement of the imaging-based classification to an established phenotype) as well as interchangeability of imaging-based classification with a comparative measure of the established phenotype are relevant and supportable claims for a multi-parametric model developed to classify phenotypes. The task we tackle in the rest of this manuscript is development of statistical methods to characterize the performance of the classifier and demonstrate agreement and/or interchangeability between imaging-derived characteristics and clinically established phenotypes. Examples of such clinical phenotype classification tasks include staging liver fibrosis or steatosis from non-invasive imaging results (rather than from pathology) or determining stability of atherosclerotic plaque from imaging findings (rather than from tissue samples). Data-driven studies intended to identify groupings within a heterogenous patient populations1921 may be extremely useful in identifying phenotypes and use case #4 provides guidance on such exploratory studies. Such clustering could be considered for this use case once clinical recognition of the phenotype is established. Furthermore, we focus the work of this manuscript on the multi-class situation, as a wealth of previous literature exists for the dichotomous (binary) classification case12,2230.

2-. Model Development

The task we tackle here is combining p QIB measurements X = (X1, X2, … Xp) into an mp-QIB classifier of K possible phenotypes 1, 2, …, K. For K = 2, Y is bivariate, and phenotype status is dichotomous. For K > 2, Y is multivariate, and phenotype status is polychotomous. In this manuscript, we focus on development of models for classification of K > 2 phenotypes, which can be either nominal or ordinal.

Y is approximated with a model f(X, β) = (f1(X, β), f2(X, β), … fK(X, β)) that is a function of X and its associated parameter β, where 0 ≤ fk(X, β) ≤ 1 and ∑fk(X, β) = 1, and fk(X, β) is the probability that the subject has phenotype k. For example, in a neural network, β includes the sets of weights and biases embedded in the input and hidden layers (see the Supplement). For classification, the maximum probability may be used to assign a subject with X = x as phenotype k = argmaxi fi(x, β). If the cost of misclassifying phenotype j as k depends on j and k, then the subject may be assigned to the phenotype that minimizes expected cost (Supplement).

The cost or loss function L(Y, f(X, β)) measures the error of approximating Y with f(X, β). For classification, L(Y, f(X, β)) is often taken to be the squared error or cross-entropy (deviance) error function31. The expected value of L(Y, f(X, β)) is the risk

R(β)=LY,fX,βdPX,Y=LY,fX,βdPX,YdPX,

where P(X, Y) = P(X)P(Y|X) is the joint distribution of X and Y, P(Y| X) is the conditional distribution of Y given X, and P(X) is the marginal distribution of X.

When f is specified, the goal is to minimize R(β) over the class of functions f(X, β), β ∈ B, which can be achieved by minimizing ∫ L(Y, f(X, β))dP(Y| X) for every X 32. In many machine learning approaches, the form of f is left unspecified and R(β) is minimized over both f and β. The model is developed by fitting f(X, β) to training data. Typically the data consist of random independent identically distributed observations {(x1, y1), … (xi, yi), …, (xn, yn)} on n subjects, where, for subject i, xi=xi1,...,xip is the set p QIB values and yi is the phenotype. To guard against overfitting, the algorithm for minimizing R(β) can be stopped early or a regularization technique imposed, among possible options (Supplement). When f is specified, the fitted model is denoted Y^=fX,β^, where β^ is the estimate of β that minimizes the sample data version of R(β). If f as well as β is estimated, then the fitted model is Y^=f^X,β^.

In general, the model output Y^=f^X,β^ need not be probabilities of classification, but instead can be a continuous numeric score to which K − 1 thresholds are applied to obtain a classification of each subject into one of the K phenotypes. For example, subjects have been classified into one of the four METAVIR liver fibrosis stages F0–1, F2, F3, and F4 based on applying three thresholds to continuous ultrasound measurement of shear wave elastography (SWE) or transient elastography (TE)33.

For classification, Y^ is typically either an assignment of X to one of the phenotypes or a set of probabilities of each phenotype. Specific to this use case, we note that while mp-QIB model inputs may be QIBs, the output of phenotype classification is not a QIB.

2.1. Model Inputs

Parameters must have biological plausibility to be candidate model inputs. However, this connection alone is insufficient, and the behavior of model inputs should also be well characterized, as stability of model inputs is necessary to inform development of a robust and reliable model. At minimum, repeatability and reproducibility of candidate model inputs should be assessed. Methods for assessing repeatability and reproducibility of a QIB are described in previous literature34,35. The more desirable state is to include fully validated QIBs as model inputs.

Highly correlated QIBs should be avoided as model inputs, as high correlations are potentially an indication that the QIBs measure the same thing and only one is needed. As a general rule, when correlation between two QIBs exceed 0.8, only one QIB should be selected for inclusion in the model.

Additional items to consider when developing models are the incremental value of each model input and whether the addition of a new parameter adds value to the model. Change in the area under the receiver operating characteristic curve (AUC is the same as the c-index when outcomes are binary), integrated discrimination improvement (IDI) and continuous net reclassification improvement (NRI) are all available metrics3639. Change in AUC with and without the inclusion of the additional parameter is perhaps the most commonly used. An improvement in AUC for a model containing the new marker provides evidence that the marker adds value; the criticism of this approach being that the incremental value of the additional marker must be quite large to result in a meaningfully larger AUC. Beyond change in AUC, Pencina et.al. propose Net Reclassification Index (NRI) and Integrated Discrimination improvement (IDI) as useful metrics to measure improvements in reclassification40. Pepe et.al. emphasize the importance of defining a minimal magnitude of improvement and testing of that rather than simply the presence of any improvement38.

2.2. Different Types of Models

The three main categories of classification approaches--classical statistical methods, machine learning approaches and deep learning neural networks—are discussed in the supplemental material.

3. Model Validation

Validation of the outputs of a developed classification model is a necessary and vital step before deployment of the model.

Internal validation is often aimed to optimize model performance and is conducted using a dataset that is not fully independent from the development dataset. Common procedures for internal validation of a model include leave-k-out cross-validation41,42 and bootstrap cross-validation43,44. A QIB developed on a training dataset will be fit to measurement error patterns and other peculiarities in the dataset as well as actual signal. Internal validation of the QIB on the same dataset on which it was trained, even if done properly, will reflect those peculiarities and thus may not be robust enough for the observed performance to reproduce in new subjects. Thus internal validation is generally not enough to establish a data-driven model45.

External validation in independent datasets is a necessary component of rigorous model assessment. The aim of external validation is to generalize and establish performance of the model in a study population different than was used for model development. External validation is performed on the locked-down model in a dataset independent of the dataset(s) used to develop the model (i.e. collected at clinical sites different from the training dataset)45.

While the external evaluations of multiparametric QIB phenotype classifiers will ultimately depend on the intended use of the model and the data available for model development, three basic performance evaluations are considered further in this manuscript:

  • Precision. Closeness of agreement of replicate classifications obtained on the same object under specified conditions46.

  • Diagnostic Accuracy. Agreement of a classification method with the true (established) phenotype, as determined by a clinical reference standard47 (which for imaging classifiers is a method other than imaging).

  • Interchangeability. Two classification methods agree such that one can be swapped for the other with negligible effect on diagnostic accuracy and clinical benefit.

Diagnostic accuracy and interchangeability are clinically meaningful claims for a multi-parametric QIB model developed to classify phenotypes. An imaging-based classification is accurate if it agrees with the established (true) phenotype. An imaging-based classification is interchangeable with a comparative classification if the two classifiers can be exchanged in clinical practice without a change in the clinical management of the patient. While bias and imprecision are important to characterize, as a clinical accuracy study of an imprecise classifier is likely to fail, the clinical interpretation of bias or imprecision for a phenotype classifier is not entirely transparent. Thus, while important to characterize, a claim that a phenotype classifier meets a specific bias or imprecision goal is not recommended to be sufficient to support clinical use of the classifier.

4. Statistical Methods for Evaluating Classifier Outputs

A phenotype classification model can be designed to classify subjects into one of the phenotypes or to output the probability of each phenotype. Methods for evaluating probabilities of having two possible phenotypes either at present or in the future are discussed in Use Case 39. In this use case, we provide a non-comprehensive review of methods for evaluating precision, diagnostic accuracy, and interchangeability of phenotype classifiers, focusing our effort on the scenario of K>2 phenotypes.

4.1. Precision

Repeatability is the precision (closeness of agreement) of repeated measures when conditions of testing are held constant46. Reproducibility is the precision of repeated measures when specific conditions of testing are varied46. For quantitative measurements, e.g., QIBs, imprecision is commonly characterized by the within-subject standard deviation (wSD) or coefficient of variation (CV) of repeated measurements46. The reproducibility coefficient (RDC) is defined as 1.962 times wSD35,48. If a serial change in a quantitative measurement is greater than its RDC, then it is likely not explainable by measurement imprecision but in fact a real change in the measurand (under stated conditions). However, for phenotype classification, SD and CV do not make sense because phenotype labels (e.g., fibrosis stage) do not express distance, that is, do not constitute an interval or ratio scale variable.

4.1.1. Variation in Categorical Data

For a phenotype classification variable S with categories i = 1,2, 3, …, K, one measure of the variation is the Gini concentration index49, gτ=1i=1Kτi2, where τ = (τ1, τ2, . . . τK) and τi = Pr(S = i). For K = 2 with τ = Pr(S = 2),

gτ=1τ21τ2=2τ1τ,

or twice the Bernoulli variance.

Clinically, g(τ) is interpretable as the probability that two realizations of S fall into different categories. In precision experiments of repeated classifications taken on the same subject, g(τ) is the probability that two repeated classifier results fall into different categories.

The minimum of g(τ) is 0 when τk = 1 for one category k and τi = 0 for all other categories ik. The maximum value of g(τ) is (K − 1)/K when τi = 1/K for all i = 1,2, … K.

The sample estimate of the Gini index, its variance, and an example of their calculation are given in the supplementary material. Two other measures of categorical variation are misclassification error50 and entropy51 (Supplement).

4.1.2. Test-Retest Precision Study

To evaluate precision on subjects, a test-retest study could be conducted in which two or more repeated measures are obtained per subject over a short period of time. Variation between the repeated measures within each subject could be characterized, then pooled across subjects to obtain an overall precision estimate. For example, for QIBs, the within-subject standard deviation (wSD) can be estimated, then pooled across subjects to obtain an overall measure of imprecision, provided wSD can be assumed to be homogeneous across subjects with different measurand values. (For many quantitative measurements, a better assumption is that the within subject coefficient of variation is homogeneous across subjects).

For phenotype classification, pooling a measure of variability across subjects is less clear but may be possible for subjects having the same phenotype. If testing conditions will be varied according to factors (e.g., imaging manufacturer, instrument within manufacturer, imaging acquisition parameter, etc.), a balanced incomplete block (BIB) design in the factor levels could be considered in which subject is the block. With a BIB design, the contribution of each factor to the total variability can be quantified. Analysis of a test-retest study for a phenotype classifier is an exciting topic for future research.

4.2. Diagnostic Accuracy

A diagnostic test result is accurate if it agrees with the accepted reference value for the measurand. For our purposes, the measurand is the phenotype and the clinical reference standard is the best available non-imaging method of assessing the established phenotype. For example, for METAVIR fibrosis stage, histological evaluation of liver biopsy is the clinical reference standard because it is considered the best available reference method (even though liver biopsy has its own limitations, including sampling error and pathologist intra- and interobserver variability in histological interpretations).

4.2.1. Binary Test Accuracy

Binary diagnostic tests are classifiers that without loss of generality render either a positive or negative test result indicating presence or absence of a condition, in this case a phenotype. The pre-test probability of the phenotype is its prevalence. The post-test probability of the phenotype is the predictive value of the test. The positive predictive value (PPV) is the probability that the subject has the phenotype given the test is positive. The negative predictive value (NPV) is the probability that the subject does not have the phenotype given the test is negative. Conversely, sensitivity (Se) is the probability of a test positive in subjects with the phenotype, and specificity (Sp) is the probability of a test negative in subjects without the phenotype. The positive and negative diagnostic likelihood ratios PLR = Se / (1 − Sp) and NLR = (1 − Se)/Sp are, by Bayes Theorem, the relative changes from pre-test to post-test in the odds of having the phenotype conferred by positive and test negative results, respectively27,28.

Binary classification models for presence of a phenotype are often based on an underlying continuous value X to which a threshold c is applied to obtain without loss of generality a positive test result when X > c and a negative test result when Xc. For example, X could be the probability of having the phenotype to which a probability threshold of ‘c=0.05 could be applied to obtain the binary decision of whether the phenotype is absent or present. The receiver operating characteristic (ROC) curve52,53 is the plot of Se(c) vs. 1 – Sp(c) for all thresholds c. A summary measure is the area under the ROC curve (AUC).

A minimal but insufficient requirement for a binary test to be clinically meaningful is that it is informative for, i.e., associated with, phenotype status. A test is informative if Se > 1 – Sp, PLR > 1, NLR < 1, PLR > NLR, PPV > 1 − NPV, or, for continuous numeric scores, AUC > 0.5. Performance metrics and methodologies for evaluating binary test accuracy are reviewed in many textbooks22,23,29,30 and papers12,2428,52,53.

4.2.2. Evaluating Multi-class Classifiers with Binary Test Accuracy Measures

Binary test accuracy measures can be used to evaluate classifiers of K > 2 phenotypes. To illustrate its use, consider the study by Ferraioli et al54 comparing real-time shear wave elastography (SWE, denoted here as S) measurements with transient elastography (TE, denoted here as T) measurements in classification of METAVIR liver fibrosis stage in 121 patients with chronic hepatitis C. The clinical reference standard for METAVIR liver fibrosis stage is histological analysis of liver biopsy (denoted here as R). SWE and TE measurements were acquired in kilopascal (kPa) units using two different ultrasound systems and thresholds were applied to the kPa measurements to create the liver fibrosis classifications.

Table 2 displays the confusion matrices of TE and SWE fibrosis stage classifications by true METAVIR fibrosis stage. For SWE, the sensitivities for fibrosis stages F0–1, F2, F3, and F4 were 87.5% (42/48), 72.7% (24/33), 84.6% (11/13), and 87.5% (21/24). For TE, the sensitivities were 89.6% (43/48), 18.8% (6/32), 53.8% (7/13), and 91.7% (22/24). Each sensitivity is accompanied by multiple specificities, one for each of the other phenotypes. For example, the sensitivity for fibrosis stage F4 is accompanied by the specificities of not being F4 when the fibrosis stages are F0–1, F2, and F3. For F4, the SWE specificities were 100% (48/48), 97.0% (32/33), and 84.6% (11/13) for F0–1, F2, and F3. If the fibrosis stage prevalences in the study are unbiased, as would be expected in a prospective study with simple random sampling of subjects, then the specificities may be pooled into an overall specificity. For SWE, the overall specificity for F4 was 96.8% (91/94 = (48 + 32 + 11)/(48 + 33 + 13)).

Table 2.

Confusion matrices of ultrasound-based shear wave elastography (SWE) and transient elastography (TE) classifications of METAVIR fibrosis stage, which was determined by histological analysis of liver biopsy (reprinted from Ferraioli et al, 2012, Table 4).

METAVIR .
SWE (kPa) Class F0–1 F2 F3 F4 Ttl

SWE ≤ 7.1 F0–1 42 7 0 0 49
7.1 < SWE ≤ 8.7 F2 4 24 0 1 29
8.7 < SWE ≤ 10.4 F3 2 1 11 2 16
SWE > 10.4 F4 0 1 2 21 24

Total 48 33 13 24 118
METAVIR .
TE (kPa) Class F0–1 F2 F3 F4 Ttl

TE ≤ 6.9 F0–1 43 19 1 1 64
6.9 < TE ≤ 8.0 F2 3 6 2 0 11
8.0 < TE ≤ 11.6 F3 2 7 7 1 17
TE > 11.6 F4 0 0 3 22 25

Total 48 32 13 24 117

Class = classification of METAVIR fibrosis stage made by the classifier

The predictive values of a classifier are the probabilities that each true phenotype agrees with the classifier result, analogous to the positive and negative predictive values of binary tests. The predictive values of SWE for F0–1, F2, F3, and F4 were 85.7% (42/49), 82.8% (24/29), 68.8% (11/16), and 87.5% (21/24). For TE, they were 67.2% (43/64), 54.5% (6/11), 41.2% (7/17), and 88.0% (22/25). Predictive values depend on the prevalences of the phenotypes, thus are unbiased only if the distribution of the true phenotypes is unbiased.

If a classifier is intended to only detect a subset of the possible phenotypes, an ‘other/none of the above’ category should be included to avoid an incomplete confusion matrix or forced misclassification. For example, if an artificial intelligence (AI) algorithm is designed to detect on chest radiograph the lung pathologies pulmonary edema (cardiomegaly), adult respiratory distress syndrome (atelectasis or pneumothorax), pneumonia (pleural effusion, consolidation, nodule), aspiration pneumonitis, or pulmonary embolism55, subjects may obviously have other lung pathologies. Without an ‘other’ category, the algorithm would call another pathology as one of the AI-detectable pathologies and the confusion matrix reporting results would be erroneous.

4.2.3. Association Between Nominal Categorical Variables with Example

A highly accurate phenotype classifier will be highly associated with the true phenotype. A general overall measure of association is the proportion of explained variation (PEV) of one variable by another56,57. For quantitative variables, PEV is the coefficient of determination or squared correlation between the variables. PEV can be used to quantify the association of two categorical variables. For the liver fibrosis example, let S = i be the SWE classifier result and R = j be the reference (true) phenotype where i, j = 1,2,3,4 index the fibrosis stages F0–1, F2, F3, and F4. Let τij = Pr(S = i, R = j) be the joint probability and nij the observed count for cell (i, j) of the confusion matrix (Table 2). The total Gini variation in reference fibrosis stage R is

V^R=gτ^=1j=14τ^j2=1481182331182131182241182=0.707

where τj = Pr(R = j) and τ^j=nij/nj is the sample estimate. The average Gini variation in R conditional on SWE result S is

E^V^R|S=i=1Iτ^i1j=1Jτ^ijτ^i2=1i=1Ij=1Jτ^ij2τ^i=0.287

where τi = Pr(S = i) and τ^i=nij/ni is its sample estimate.

The proportion of Gini variation in R explained by S is therefore

PEVgτ^=V^RE^V^R|SV^TR=0.7030.2870.703=0.592

In the Supplemental Material, the proportion of variation in R explained by S is also calculated for entropy and classification error for this example. While a highly accurate phenotype classifier will be highly associated with the true phenotype, we caution that the reverse need not be true, although that would be strange to observe.

4.2.4. Association Between Ordinal Variables

A symmetrical measure of association between ordinal classifications S and R with I and J categories is Gamma56,58. Like correlation, γ ranges from −1 to 1. A pair of subjects is considered concordant if the subject ranked higher on S also ranks higher on R. A pair of subjects is considered discordant if the subject ranked higher on S ranks lower on R. The calculation of Gamma for the liver fibrosis example is given in the supplemental material. For the SWE classification data γ^=0.9429; for the TE classification data, γ^=0.9041. Alternatively, for I > 2 and J > 2 ordinal categories, Obuchowski (2005)59 proposed an AUC-like measure (Matlab code provided in Supplemental material).

4.2.5. Bias

Bias in a classifier can contribute to its inaccuracy. To evaluate bias, express the classifier result S = i as the multinomial variable S = (S1, S2, … SK) with Si = 1 and Sk = 0 for ki, and likewise express the reference (true) phenotype as R = (R1, R2, … RK). The marginal distributions of S and R are τS = (τ1•, τ2•, … τK) and τR = (τ•1, τ•2, … τK). The multivariate bias of S for R is the difference D = τSτR. The null hypothesis of no bias is H0: D = 0, which is the hypothesis of marginal homogeneity of row and column proportions. For dichotomous data, the McNemar test statistic of H0 is

Z=n12n21n12+n211/2

where nij is the observed count in cell (i, j) the confusion matrix. For K by K tables, marginal homogeneity can be evaluated with Stuart’s score test56,60 which when H0 is true asymptotically follows a standard normal distribution (mean 0, variance 1). For illustration of these tests with the SWE data, see the Supplemental Material. If the hypothesis test rejects H0, then one can conclude that S is biased for R. If the hypothesis test does not reject H0, then one cannot conclude that S is unbiased because the power to reject H0 may have been low due to limited sample size. Rejection of H0 does not necessarily mean that the bias of S for R is clinically significant. A multivariate confidence interval on D would be helpful in assessing clinical significance of the bias by comparing the interval with acceptable bias thresholds for each phenotyope category (Supplement).

4.2.6. Symmetry

A classifier and the reference phenotype can have homogeneous marginal proportions, yet asymmetric joint probabilities of classification. Asymmetry is a sign that the classifier is locally biased for some phenotypes. For example, in the liver fibrosis data (Table 2), TE called 19 cases of fibrosis stage F2 as F0–1, but only 3 F0–1 cases as F2. A description of the Bowker’s Pearson X2 test for symmetry61, as well as illustration with the SWE data, is given in the Supplemental Material.

4.2.7. Association and Agreement Models

Association models for categorical variables have parameters such as odd ratios that quantify association between the variables. Logistic and multinomial regression models designate one variable as the outcome and the other variables as predictors. Log linear models are agnostic, i.e., do not designate variables as either outcome or predictor.

Agreement models are less commonly used, but include parameters that quantify agreement between variables. A review of agreement models is Agresti (1992)62. In the Supplemental Material, we illustrate various association and agreement models and metrics using the liver fibrosis SWE and TE data.

4.3. Interchangeability

For quantitative measurements, interchangeability assessment is reviewed in the Supplemental Material. For classifiers, interchangeability assessment has received relatively little attention. Obuchowski et al (2014)63 discuss the concept of interchangeability of imaging tests.

When a reference standard for determining the true phenotype is unavailable, the accuracy of a classifier cannot be evaluated. Unfortunately, the literature is replete with examples in which classification accuracy is nonetheless “evaluated” based on agreement with an imperfect reference standard. When the classifier and imperfect reference standard are conditionally independent given the true phenotype, classifier accuracy will tend to be underestimated with this evaluation. However, serious overstatements of accuracy can occur when the classifier and the imperfect reference standard are positively correlated. Common reasons for positive correlation are (1) the classifier was trained to agree with the imperfect reference standard, and (2) the classifier and imperfect reference standard are based on the same or similar measurement procedure(s), that is, examination of the same medium or milieu (e.g., image or biological sample), measurement of the same or similar analyte, application of the same or similar technology by which measurement(s) are taken, and/or application of the same or similar calculation algorithm for obtaining the final classification result. An example of (1) is a machine learning classifier trained to agree with reader interpretations of images that is validated for accuracy in a separate study to classifier agreement with a consensus of reader interpretations. An example of (2) is evaluation of the “accuracy” of a nucleic acid amplification test (NAAT) for chlamydia trachomatis based on agreement with the consensus of two comparator NAAT assays according to the patient infection status algorithm64. Pepe (2003) provides numerical examples of the types of bias in diagnostic test accuracy that are introduced depending on the relationship between a test and an imperfect reference standard22.

When a reference standard is unavailable for determining the phenotype, the accuracy of a classifier cannot be estimated. In the absence of a reference standard, a classifier variable S could be evaluated for interchangeability with a clinically accepted comparator classifier T based on their agreement. A classifier is interchangeable with a comparator if one can be swapped for the other with negligible change in clinical benefit. Interchangeability would be a meaningful claim to have for a classifier if it had advantages over the comparator, e.g., less invasive, less time-consuming, less costly, more automated, or more reproducible.

The effect on clinical benefit of swapping S for T should be negligible when (a) S agrees with T as much as T agrees with itself in replicate testing, and (b) the distribution of disagreements between S and T is symmetric, as is assumed to be true for replicate values of T. Criterion (a) is the concept of individual equivalence65,66. Obuchowski (2001)67 evaluated individual equivalence of electronic medical images (e.g., full-field digital mammography) with screen film, comparing the number of cases with ‘equivalent’ reader diagnoses on electronic images and film to the number of cases with ‘equivalent’ diagnoses when the readers are just using film.

Symmetry is related to the concept of exchangeability. S and T are exchangeable if their labels can be permuted without altering their joint distribution. Symmetry of S and T cannot be shown with statistical significance because estimates of cell probabilities are uncertain in finite samples. For hypothesis testing, non-inferiority margin(s) in asymmetry of the cell probabilities could be pre-specified and tested for statistical significance.

We contend that agreement and symmetry are the hallmarks of interchangeability. Agreement is not the same as association. As mentioned, two categorical variables can be highly associated yet agree poorly because their joint distribution is asymmetric. However, strong association together with symmetry implies high agreement. Thus, interchangeability of two classifiers can be assessed by examining their association and symmetry.

5. A Case Study: Atherosclerotic Plaque

In this section we provide a real-world example of a phenotype classification task, taking the example from an understanding of the clinical context to model development, and finally model evaluation. Figure 1 provides a pictorial overview of model development process.

Figure 1:

Figure 1:

An overview of the atherosclerotic plaque classification task. A) Standardized CTA images are acquired in accordance with the existing QIBA profile. B) Images are used to classify tissue type and location at each cross section along the vessel wall. Vessel morphology is validated against histological truth. C) Cross sectional views along the vessel depict tissue type and location as color overlays to the vessel wall. D) Area and location of tissue at each cross section is used as input into a convolutional neural network to generate a phenotype classification (subclinical, stable, unstable) at each vessel cross section. In this way, variations in raw images are filtered out, and focus remains on validated tissue characteristics. Phenotype classification is compared against agreement with an expert pathologist blinded to imaging results.

5.1. Clinical Context

Despite advances in therapy, reduction in myocardial infarction or death remains elusive. Cardiovascular disease is the most common cause of death and disability in the world, mainly by myocardial infarction and ischemic stroke resulting from unstable atherosclerosis68. A clear need exists for more precise patient categorization. Despite discoveries of new predictive plasma biomarkers69, routine diagnostic methods for identification of individuals and lesions at high risk for atherothrombosis in coronary or extracranial arteries are still lacking70.

5.2. Classification Task

In this example, we aim to categorize plaque stability phenotype from computed tomography angiography (CTA) images. The plaque phenotype classification system approved by the American Heart Association (AHA) to inform practicing clinicians on the current presentation of the plaque is tied to levels of severity regarding propensity to rupture. Stary’s initial system71 has been updated and refine in recent years72,73. For this application, the granularity of the AHA system was reduced to strike an optimal balance between feature visibility in radiology and the clinical utility of the classification. We aimed to maximize the value of in vivo radiology relative to clinical utility by identifying what a pathologist would say were they to have the tissue at microscopy, but doing so without tissue. From a mathematical point of view, phenotype classification outputs are polychotomous multiclass outcomes (minimal disease, stable plaque, and unstable plaque) which are not constrained physiologically for a certain order of progression beyond the fact that all disease starts as minimal.

5.3. Model Development

The task is classification of atherosclerotic plaques as minimal disease, stable plaque, or unstable plaque. The labelling of plaque type at microscopy sections paired with computed tomography angiography (CTA) serves as the reference standard. Specifically, CTA tissue classification (validated against microscopy sections with outlined tissue type) is the validated model input (Figure 1A-C), while correctness of classification (that is, model output) is assessed relative to a consensus panel of pathologists. Both are blinded to imaging results.

When deep neural networks are trained, the initial layers learn to detect general shapes, edges, corners, etc. The deeper layers learn fine-grained features of the data. While using transfer learning the initial layers are frozen as to not re-learn the task of general shape/edge/corner detection, but the later layers are unfrozen so that the model can update its learning with the current data. The number of frozen layers is set as a hyperparameter. Given the clinical implications of certain types of misclassifications, a non-constant cost function is most appropriate.

5.4. Model Inputs

CTA has been identified as the most effective tool for assessing ischemic heart disease in the prospective Evaluation of Integrated Cardiac Imaging in Ischemic Heart Disease (EVINCI) comparison74. Trials such as the Scottish Computed Tomography of the HEART Trial (SCOT-HEART)75, and ISCHEMIA76 are increasingly showing the sensitivity of the modality. Although qualitative CTA remains the prevalent clinical modality, quantitative CTA analysis offers great potential for biomarker identification and development. Moreover, there is increasing evidence linking quantitative CTA measurements to molecular phenotype77. In this example, quantitative measurements are made with ElucidVivo (Boston, MA) software78.

This example benefits from knowledge of the associated disease process. Specifically, it is well known that if atherogenic factors dominate over atheroprotective factors, plaques grow in size, and three basic tissue types start to dominate: extra-cellular matrix (ECM) and/or fibrosis may dominate; vascular smooth muscle cells may become osteogenic which in turn promotes calcification, first as micro calcification and later as macro-calcification; or lipid-rich necrotic core (LRNC) and intra-plaque hemorrhage (IPH) may dominate. Plaques largely composed of ECM and/or fibrosis or macro-calcification are generally considered stable; micro-calcification and/or LRNC/IPH are generally considered to be unstable.

QIBA has published a profile for quantitative biomarkers in atherosclerosis79. Quantitative measurements included in the profile include maximum wall thickness, lumen area, wall area, plaque burden, lipid-rich necrotic core area, and calcified area. The QIBA profile outlines procedures for taking measurements within a prescribed anatomical target comprising of one or more vessels, and at perpendicular cross-sections along the centerline of each vessel (Figure 1B). Each cross-section thereby presents as a roughly circular lumen area (representing the blood channel) and an annular wall area (presenting the vessel wall, including plaque with its constituent tissues). These quantitative CTA measurements are used as the basis for input to the classification model. However, we expand upon the scalar measurements outlined in the QIBA profile to include not only magnitude of tissue composition but also location and relationship between tissues at a given location. Thus, quantitative information about the location and type of tissue, rather than the raw CTA images, are used as model inputs (Figure 1C). This approach was chosen rather than a combination of scalar QIBs (such as lumen area, wall thickness, calcified area, etc.) because not only the amounts of tissues, but their spatial relationship, is significant. Additionally, an understanding of the underlying pathological processes allowed focus to remain only on information relevant to that process. Characterization of the tissues individually constitutes a single measurand scenario where identified tissues are QIBs with reference truth by histology79.

The phenotype is determined from a combination of the individual validated single QIBs from the QIBA CTA profile, not by any single one (Figure 1D). The raw CTA images themselves are not used as model inputs, as raw images are not validated, and include excess variability not tied to the real phenotype that can cause overfitting in models. Rather, we use the tissue characterization first, and feed color-coded images that have been normalized with respect to the dimensions into the model.

Candidate methods need to mitigate drawbacks of machine learning to keep the benefit of the non-linearity of the network. One drawback is how relatively few samples can be relied on, given the relative difficulty of obtaining annotated samples in medicine as compared to other applications of computer vision, and also, how to ensure there is a mechanistic rationale based in biology rather than the algorithm being a black box.

5.5. Model Validation

Model outputs will be assessed relative to agreement with an expert pathologist with results pending at time of publication. In accordance with good model development practices, model evaluation is performed on a separate test dataset from that used to develop and train the model. Correctness of classification by the developed model is assessed relative to a consensus panel of pathologists blinded to imaging results. Agreement of classification by the developed model (i.e., classifying as v=1 (minimal), 2 (stable), and 3 (unstable)) with the corresponding truth states (t=1 (minimal), 2 (stable), and 3 (unstable)) is calculated via the weighted kappa statistic that indicates the proportion of agreement between the assessment categories of two classification results, beyond that expected by chance, penalizing disagreements in terms of their seriousness. The weighted kappa, quantified as κ = (pope)/(1 − pe), compares the observed agreement po=t=13v=13wtvptv with the chance expected agreement pe=t=13v=13wtvptpv, where ptv, pt, pv are the joint and the marginal proportions of the corresponding cross-classification table80. The weights located at the diagonal represent agreement and thus equal 1. Off diagonal weights indicate the seriousness of disagreement and for the quadratic weighting scheme are given by wtv = 1 − |tv|2/4.

The supportable claim is the ability of the model to classify atherosclerotic plaques into one of three phenotypes: “minimal” (which includes normal), “stable”, or “unstable” with a high level of agreement relative to expert assessed histopathology. This could be viewed as a cross-section diagnostic biomarker per the NIH BEST definitions.

6. Discussion

We have provided a discussion on the utility of combining multiple quantitative imaging biomarkers to perform a phenotype classification task and proposed methodology through which it is possible to assess the performance of such a classifier and demonstrate accuracy or interchangeability of the classification task. The methods proposed here build upon existing methodology for univariate quantitative imaging biomarkers and attempt to extend those methods to multiparametric classification tasks. We also reviewed statistical methods for developing a model for a phenotype classification task. Finally, we provided a real-world, end-to-end example in which an imaging-based deep neural network was developed and validated to classify atherosclerotic plaques as minimal, stable, or unstable based on CTA imaging.

6.1. Challenges unique to this use case

Accuracy of a measurement is traditionally assessed by comparison to an objectively measurable reference value. In phenotype classification, the clinically established phenotype is accepted as the reference value to which the imaging-derived phenotype is compared. This may present a limitation to the claims that can be made for this phenotype classifiers because for many disease conditions, the phenotype cannot be established unequivocally, as a clinical reference standard for determining the phenotype is often unavailable. For example, majority consensus interpretation of images is sometimes used as the reference in radiological applications. However, majority consensus is typically imperfect and thus not clinical proof of the phenotype81. Moreover, an algorithm for a phenotype classifier trained to agree with reader image interpretations would result in the classifier and majority consensus being positively correlated. That is, the classifier would tend to agree with misclassification errors made by majority consensus, resulting in classifier accuracy being overstated. Therefore, for classifier accuracy evaluation, we limited our profile development efforts to the subset of phenotypes for which a reference standard/best available method exists and is based upon something other than radiological imaging (e.g., histopathology or genomic expression). It is recognized that radiological imaging may not always be able to capture a phenotype with the same granularity as the reference method; thus sub-classification or grouping of an established phenotype into coarser metrics may be necessary for comparison to the imaging-based method. When a reference standard does not exist and accuracy cannot be evaluated, a classifier may be assessed for interchangeability with an imperfect reference or a commonly used comparator classifier by evaluating the two classification algorithms for association and symmetry. However, more research is needed on interchangeability assessment of classifiers.

Because classifications are not quantitative, the within-subject standard deviation (wSD) cannot be used to characterize the imprecision (measurement variability) of the phenotype classification task. To quantify categorical variation, we considered the Gini index, entropy, and classification error, which are valid for nominal variables but to do not exploit the ordinal nature of ordinal variables.

Model-based phenotype classification is a many-to-one function of input QIBs. This makes quantification of model imprecision even more difficult, as the phenotype class depends on imprecision in QIB input variables and their correlations. Imprecision of phenotype class (its distribution in replicate classification results) can be approximated analytically from imprecision of QIB inputs using the multivariate delta method82. The delta method approximation is accurate in large samples but may be poor in small samples and may be prohibitive to compute for complex, non-linear models. Simulation may be necessary. Finally, imprecision of QIB inputs (e.g., wSDs) might not be constant, e.g., might be proportional to the measurand reference value. Thus, imprecision of phenotype class may be larger or smaller depending on the input values that rendered the same output class. Further study is warranted in how best to combine imprecision of individual model inputs to quantify imprecision of the classification model output. While we have focused on the model output being the phenotype assignment, it could also be a vector of probabilities of each phenotype with assignment based on minimizing expected cost, or a continuous score to which thresholds are applied to determine assignment. We recognize that the same score arising from different model inputs may have different imprecision. Whether (and how) information on score imprecision should be retained to help inform clinical decision making likely depends upon the individual use case. Imprecision quantification of phenotype classification may require special considerations and statistical methods.

An area of continued discussion is the needed level of validation for model input variables. While we recognize that fully validated QIBs with clear biological connection are the ideal situation, we stop short in this manuscript of making that unilateral recommendation. Rather we state that model inputs should be well-characterized and stability of model inputs is necessary to inform development of a robust and reliable model. This position is intended to balance model performance against overly burdensome development activities.

6.2. Gaps

Evaluation of phenotype classifiers is a developing and evolving space. The need for biological explain-ability or interpretability of models based on machine learning influences both the technical design of proposed methods and their validation as a phenotype classifier, but is absolutely not unique to this use case83. Especially when classification models are developed using deep learning networks, use of objectively validated inputs rather than raw images coupled with understanding the characteristics of model decisions is important to ensuring confidence in the model.

We have assumed that the phenotype classifications are mutually exclusive and exhaustive. In our example, the METAVIR fibrosis stages are mutually exclusive and exhaustive for the patient population. However, in some clinical settings a subject can simultaneously have more than one phenotype. For example, a head CT scan may indicate presence of one or more conditions such as intracranial hemorrhage, mass effect, midline shift, or cranial fracture84. In this setting, many combinations of phenotypes are possible. Evaluating the accuracy of a classifier to correctly identify every particular combination can be a daunting task and may require statistical adjustment for multiplicity or models to smooth estimates.

We assumed data were cross-sectional and thus did not consider disease stage transitions over time. Over time subjects can progress from an early disease stage to a later disease stage. Subjects can also regress back to an earlier stage from a later stage. For example, stage 2 cervical intraepithelial neoplasia, can regress back to normal especially in young women85. Multistate models of event history data can model the probabilities of disease transitions over time, which can be helpful for determining an appropriate screening interval. The models can be extended to incorporate misclassification error86.

Clinical benefit, also called clinical utility87 or patient outcome efficacy88, can be defined as the effect a phenotype classifier has on efficacy outcome(s) when it is used in the intended use population per instructions for use. In clinical outcome trials of diagnostic tests for statistical efficiency, the treatments under study are randomly allocated to the subset of subjects in whom the test result would change the treatment decision if the tests were used in practice89,90. While we did not consider decision analysis, it can be used to quantify the net benefit of a diagnostic test based on the clinical consequences of testing and diagnostic error91.

In evaluating classifier accuracy, agreement models were considered in the Supplementary Material. These models contain parameters specific for assessing agreement (as opposed to association). As more experience is gained with the parameters of agreement models, they are expected to provide further insights into the evaluation of classifier accuracy and interchangeability.

Phenotype classifiers S and T were considered interchangeable if swapping one for the other in a clinical practice setting would result in a negligible change on clinical benefit to patients. The interchangeability assessments of association and symmetry presented in Section 4.3 and in the supplementary material are novel in QIB evaluation. Association and symmetry were considered for assessing interchangeability because together they imply high agreement and because symmetry is required for the result of one classifier to behave as if it were a repeated measure from the other classifier. More experience is needed before these assessments can be recommended for wide adoption.

Many if not most machine learning models do not provide the uncertainty of model output. For example, a phenotype classification model that reports the probabilities of having the phenotypes typically does not also report confidence intervals on these probabilities. How best to address uncertainty of the classification output remains an open area of consideration. We believe that communicating the uncertainty of a classifier result to the user facilitates proper use of results in clinical decision making.

How to evaluate if model-estimated probabilities of phenotypes are well-calibrated with actual probabilities is discussed in the overview7 and Use Case 3 (Huang et al)9. Well-calibrated probabilities can be useful for individualized clinical decision support. For example, the probability of a phenotype may be much higher than the probabilities of the other possible phenotypes and may exceed a threshold recommended for taking clinical action. Model-estimated probabilities of phenotypes can be regarded as updates to the prevalences of the phenotypes given the model inputs. If the model was trained on dataset(s) with phenotype prevalences that are distorted relative to those in the population on whom the model will be used, then the model-estimated probabilities will not be well-calibrated to the actual probabilities and will require adjustment before they can be clinically useful in that population

7. Conclusions

We have outlined an approach and statistical methodology for development of a phenotype classification task from a set of multiparametric QIBs. We have outlined performance metrics to consider and the diagnostic accuracy and interchangeability claims we believe are relevant and supportable for a phenotype classification task informed by multiple QIB model inputs. We discuss various approaches to model development and present options for statistical methodology that can be useful in assessment of the developed model. We intend this manuscript helps to further and advance the robust and promising science of multiparametric biomarker development.

Supplementary Material

1

Footnotes

Publisher's Disclaimer: This is a PDF file of an unedited manuscript that has been accepted for publication. As a service to our customers we are providing this early version of the manuscript. The manuscript will undergo copyediting, typesetting, and review of the resulting proof before it is published in its final form. Please note that during the production process errors may be discovered which could affect the content, and all legal disclaimers that apply to the journal pertain.

Declaration of interests

The authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper.

Contributor Information

Jana G Delfino, Center for Devices and Radiological Health, US Food and Drug Administration. Silver Spring, MD.

Gene A Pennello, Center for Devices and Radiological Health, US Food and Drug Administration, Division of Imaging, Diagnostic and Software Reliability, Office of Science and Engineering Laboratories, Center for Devices and Radiological Health, US Food and Drug Administration, 10903 New Hampshire Avenue, Silver Spring, MD 20993.

Huiman X Barnhart, Department of Biostatistics and Bioinformatics, Duke University Durham, NC, USA.

Andrew J Buckler, Elucid Bioimaging, Inc. Boston, MA.

Xiaofeng Wang, Department of Quantitative Health Sciences, Lerner Research Institute, Cleveland Clinic, Cleveland, OH, USA.

Erich P Huang, Biometric Research Program, Division of Cancer Treatment and Diagnosis – National Cancer Institute, National Institutes of Health, Bethesda, MD.

Dave L Raunig, Data Science Institute, Statistical and Quantitative Sciences, Takeda Pharmaceuticals America Inc, Lexington, MA, USA.

Alexander R. Guimaraes, Department of Diagnostic Radiology, Oregon Health & Sciences University, Portland, OR, USA

Timothy J Hall, Department of Medical Physics, University of Wisconsin, Madison, WI, USA.

Nandita N.M. deSouza, Division of Radiotherapy and Imaging, The Institute of Cancer Research and Royal Marsden NHS Foundation Trust, London, United Kingdom and European Imaging Biomarkers Alliance (EIBALL), European Society of Radiology (ESR), Am Gestade 1, Vienna, Austria.

Nancy Obuchowski, Department of Quantitative Health Sciences, Lerner Research Institute Cleveland Clinic, 9500 Euclid Ave/JJN3, Cleveland, OH 44195, USA.

8 References

  • 1.Sullivan DC, Obuchowski NA, Kessler LG, et al. Metrology Standards for Quantitative Imaging Biomarkers [published online ahead of print 2015/08/13]. Radiology 2015;277(3):813–825. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2.Kessler LG, Barnhart HX, Buckler AJ, et al. The emerging science of quantitative imaging biomarkers terminology and definitions for scientific studies and regulatory submissions [published online ahead of print 2014/06/13]. Stat Methods Med Res 2015;24(1):9–26. [DOI] [PubMed] [Google Scholar]
  • 3.Raunig DL, McShane LM, Pennello G, et al. Quantitative imaging biomarkers: a review of statistical methods for technical performance assessment [published online ahead of print 2014/06/13]. Stat Methods Med Res 2015;24(1):27–67. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.Obuchowski NA, Reeves AP, Huang EP, et al. Quantitative imaging biomarkers: a review of statistical methods for computer algorithm comparisons [published online ahead of print 2014/06/13]. Stat Methods Med Res 2015;24(1):68–106. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Obuchowski NA, Barnhart HX, Buckler AJ, et al. Statistical issues in the comparison of quantitative imaging biomarker algorithms using pulmonary nodule volume as an example [published online ahead of print 2014/06/13]. Stat Methods Med Res 2015;24(1):107–140. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Huang EP, Wang XF, Choudhury KR, et al. Meta-analysis of the technical performance of an imaging procedure: guidelines and statistical methodology [published online ahead of print 2014/05/30]. Stat Methods Med Res 2015;24(1):141–174. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.Obuchowski N, Huang E, deSouza N, et al. A Framework for Evaluating the Technical Performance of Multiparameter Quantitative Imaging Biomarkers (mp-QIBs):. Academic Radiology 2022;submitted. [DOI] [PMC free article] [PubMed]
  • 8.Raunig D, Delfino JG, Pennello G, et al. Multiparametric Quantitative Imaging Biomarker as a Multivariate Descriptor of Health. Academic Radiology 2022;submitted. [DOI] [PMC free article] [PubMed]
  • 9.Huang E, Pennello G, deSouza N, et al. A Roadmap for Developing and Evaluating Quantitative Imaging Biomarker-Based Models for Risk Prediction. Academic Radiology 2022;submitted.
  • 10.Wang X, Pennello G, deSouza N, et al. Multiparametric Data-driven Imaging Markers: Guidelines for Development, Application and Reporting of Model Outputs in Radiomics. Academic Radiology 2022;submitted. [DOI] [PMC free article] [PubMed]
  • 11.Hoehndorf R, Schofield PN, Gkoutos GV. Analysis of the human diseasome using phenotype similarity between common, genetic, and infectious diseases [published online ahead of print 2015/06/09]. Sci Rep 2015;5:10888. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Swets JA. Measuring the accuracy of diagnostic systems [published online ahead of print 1988/06/03]. Science 1988;240(4857):1285–1293. [DOI] [PubMed] [Google Scholar]
  • 13.Amarenco P, Bogousslavsky J, Caplan LR, Donnan GA, Hennerici MG. Classification of stroke subtypes [published online ahead of print 2009/04/04]. Cerebrovasc Dis 2009;27(5):493–501. [DOI] [PubMed] [Google Scholar]
  • 14.Kinner S, Reeder S, Yokoo T. Quantitative Imaging Biomarkers of NAFLD [published online ahead of print 2016/02/06]. Dig Dis Sci 2016;61(5):1337–1347. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Cascianelli S, Molineris I, Isella C, Masseroli M, Medico E. Machine learning for RNA sequencing-based intrinsic subtyping of breast cancer [published online ahead of print 2020/08/23]. Sci Rep 2020;10(1):14071. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Goodman ZD. Grading and staging systems for inflammation and fibrosis in chronic liver diseases [published online ahead of print 2007/08/19]. J Hepatol 2007;47(4):598–607. [DOI] [PubMed] [Google Scholar]
  • 17.Waxman AG, Chelmow D, Darragh TM, Lawson H, Moscicki AB. Revised terminology for cervical histopathology and its implications for management of high-grade squamous intraepithelial lesions of the cervix [published online ahead of print 2012/11/22]. Obstet Gynecol 2012;120(6):1465–1471. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Armato SG 3rd, Huisman H, Drukker K, et al. PROSTATEx Challenges for computerized classification of prostate lesions from multiparametric magnetic resonance images [published online ahead of print 2019/03/07]. J Med Imaging (Bellingham) 2018;5(4):044501. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.Calfee CS, Delucchi K, Parsons PE, Thompson BT, Ware LB, Matthay MA. Subphenotypes in acute respiratory distress syndrome: latent class analysis of data from two randomised controlled trials. The Lancet Respiratory Medicine 2014;2(8):611–620. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.Sakr L, Small D, Kasymjanova G, Suissa S, Ernst P. Phenotypic heterogeneity of potentially curable non-small-cell lung cancer: cohort study with cluster analysis [published online ahead of print 2015/04/23]. J Thorac Oncol 2015;10(5):754–761. [DOI] [PubMed] [Google Scholar]
  • 21.Wu J, Cui Y, Sun X, et al. Unsupervised Clustering of Quantitative Image Phenotypes Reveals Breast Cancer Subtypes with Distinct Prognoses and Molecular Pathways [published online ahead of print 2017/01/12]. Clin Cancer Res 2017;23(13):3334–3342. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Pepe MS. The Statistical Evaluation of Medical Tests for Classification and Prediction Oxford; 2003. [Google Scholar]
  • 23.Zhou ZH ON, McClish DK. Statistical Methods in Diagnostic Medicine 2nd Edition ed: Wiley; 2010. [Google Scholar]
  • 24.Altman DG, Bland JM. Diagnostic tests. 1: Sensitivity and specificity [published online ahead of print 1994/06/11]. BMJ 1994;308(6943):1552. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25.Altman DG, Bland JM. Diagnostic tests 2: Predictive values [published online ahead of print 1994/07/09]. BMJ 1994;309(6947):102. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26.Altman DG, Bland JM. Diagnostic tests 3: receiver operating characteristic plots [published online ahead of print 1994/07/16]. BMJ 1994;309(6948):188. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27.Deeks JJ, Altman DG. Diagnostic tests 4: likelihood ratios [published online ahead of print 2004/07/20]. BMJ 2004;329(7458):168–169. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28.Simel DL, Samsa GP, Matchar DB. Likelihood ratios with confidence: Sample size estimation for diagnostic test studies. Journal of Clinical Epidemiology 1991;44(8):763–770. [DOI] [PubMed] [Google Scholar]
  • 29.ZKLABAO-MLR HE. Statistical Evaluation of Diagnostic Performance: Topics in ROC Analysis Chapman & Hall; 2012. [Google Scholar]
  • 30.Krzanowski W, Hand D. ROC Curves for Continuous Data Chapman and Hall; 2009. [Google Scholar]
  • 31.Hastie T, Tibshirani R and Friedman J The elements of statistical learning: data mining, inference, and prediction Springer Science & Business Media; 2009. [Google Scholar]
  • 32.Anderson T. An Introduction to Multivariate Statistical Analysis 3rd edition ed 2003.
  • 33.Barr RG, Ferraioli G, Palmeri ML, et al. Elastography Assessment of Liver Fibrosis: Society of Radiologists in Ultrasound Consensus Conference Statement [published online ahead of print 2015/06/17]. Radiology 2015;276(3):845–861. [DOI] [PubMed] [Google Scholar]
  • 34.Bachtiar V, Kelly MD, Wilman HR, et al. Repeatability and reproducibility of multiparametric magnetic resonance imaging of the liver [published online ahead of print 2019/04/11]. PLoS One 2019;14(4):e0214921. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 35.Standardization IOf. Guidance for the use of repeatability, reproducibility and trueness estimates in measurement uncertainty estimation. In. ISO Standard No 21748:2017(E) Geneva, Switzerland: 2017. [Google Scholar]
  • 36.Kerr KF, McClelland RL, Brown ER, Lumley T. Evaluating the incremental value of new biomarkers with integrated discrimination improvement [published online ahead of print 2011/06/16]. Am J Epidemiol 2011;174(3):364–374. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 37.Kerr KF, Bansal A, Pepe MS. Further insight into the incremental value of new markers: the interpretation of performance measures and the importance of clinical context [published online ahead of print 2012/08/10]. Am J Epidemiol 2012;176(6):482–487. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 38.Pepe MS, Kerr KF, Longton G, Wang Z. Testing for improvement in prediction model performance [published online ahead of print 2013/01/09]. Stat Med 2013;32(9):1467–1482. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 39.Vickers AJ, Cronin AM, Begg CB. One statistical test is sufficient for assessing new predictive markers [published online ahead of print 2011/02/01]. BMC Med Res Methodol 2011;11:13. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 40.Pencina MJ, D’Agostino RB Sr., D’Agostino RB Jr., Vasan RS. Evaluating the added predictive ability of a new marker: from area under the ROC curve to reclassification and beyond [published online ahead of print 2007/06/15]. Stat Med 2008;27(2):157–172; discussion 207–112. [DOI] [PubMed] [Google Scholar]
  • 41.Ambroise C, McLachlan GJ. Selection bias in gene extraction on the basis of microarray gene-expression data [published online ahead of print 2002/05/02]. Proc Natl Acad Sci U S A 2002;99(10):6562–6566. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 42.Simon R, Radmacher MD, Dobbin K, McShane LM. Pitfalls in the use of DNA microarray data for diagnostic and prognostic classification [published online ahead of print 2003/01/02]. J Natl Cancer Inst 2003;95(1):14–18. [DOI] [PubMed] [Google Scholar]
  • 43.Efron B, Tibshirani R. Improvements on Cross-Validation: The .632+ Bootstrap Method. Journal of the American Statistical Association 1997;92(438). [Google Scholar]
  • 44.Harrell FELK, and Mark DB. Tutorial in biostatistics: multivariable prognostic models: issues in developing models, evaluating assumptions and adequacy, and measuring and reducing errors. Statistics in Medicine 1996;15:361–387. [DOI] [PubMed] [Google Scholar]
  • 45.Altman DG, Royston P. What do we mean by validating a prognostic model? Statistics in Medicine 2000;19(4):453–473. [DOI] [PubMed] [Google Scholar]
  • 46.CLSI Harmonized Terminology Database In. doi: https://htd.clsi.org/
  • 47.Bossuyt PM, Reitsma JB, Bruns DE, et al. STARD 2015: An Updated List of Essential Items for Reporting Diagnostic Accuracy Studies [published online ahead of print 2015/10/29]. Radiology 2015;277(3):826–832. [DOI] [PubMed] [Google Scholar]
  • 48.Raunig DL ML, Pennello G, Gatsonis C, Carson PL, Voyvodic JT, Wahl RL, Kurland BF, Schwarz AJ, Gönen M, Zahlmann G, Kondratovich MV, O’Donnell K, Petrick N, Cole PE, Garra B, Sullivan DC;. Quantitative imaging biomarkers: a review of statistical methods for technical performance assessment. Stat Methods Med Res 2015;24(1):27–67. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 49.Light R, Margolin B. An Analysis of Variance for Categorical Data. Journal of the American Statistical Association 1971;66:534–544. [Google Scholar]
  • 50.MittlbÖCk M, Schemper M. Explained Variation for Logistic Regression. Statistics in Medicine 1996;15(19):1987–1997. [DOI] [PubMed] [Google Scholar]
  • 51.Haberman SJ. Analysis of Dispersion of Multinomial Responses. Journal of the American Statistical Association 1982;77(379). [Google Scholar]
  • 52.Zweig MH CGA-EiCCAP. Receiver-operating characteristic (ROC) plots: a fundamental evaluation tool in clinical medicine. Clin Chem 1993;39(4):561–577. [PubMed] [Google Scholar]
  • 53.Obuchowski NA. ROC analysis [published online ahead of print 2005/01/27]. AJR Am J Roentgenol 2005;184(2):364–372. [DOI] [PubMed] [Google Scholar]
  • 54.Ferraioli G, Tinelli C, Zicchetti M, et al. Reproducibility of real-time shear wave elastography in the evaluation of liver elasticity [published online ahead of print 2012/07/04]. Eur J Radiol 2012;81(11):3102–3106. [DOI] [PubMed] [Google Scholar]
  • 55.Khan AN, Al-Jahdali H, Al-Ghanem S, Gouda A. Reading chest radiographs in the critically ill (Part II): Radiography of lung pathologies common in the ICU patient [published online ahead of print 2009/07/31]. Ann Thorac Med 2009;4(3):149–157. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 56.Agresti A. Categorical Data Analysis 3rd Edition ed 2013.
  • 57.Agresti A. Applying R2-type measures to ordered categorical data. Technometrics 1986;28:133–138. [Google Scholar]
  • 58.Goodman LA, Kruskal WH. Measures of Association for Cross Classifications*. Journal of the American Statistical Association 1954;49(268):732–764. [Google Scholar]
  • 59.Obuchowski NA. Estimating and comparing diagnostic tests’ accuracy when the gold standard is not binary [published online ahead of print 2005/08/16]. Acad Radiol 2005;12(9):1198–1204. [DOI] [PubMed] [Google Scholar]
  • 60.Stuart A A Test for Homogeneity of the Marginal Distributions in a Two-Way Classification. Biometrika 1955;42(3–4):412–416. [Google Scholar]
  • 61.Bowker AH. A test for symmetry in contingency tables [published online ahead of print 1948/12/01]. J Am Stat Assoc 1948;43(244):572–574. [DOI] [PubMed] [Google Scholar]
  • 62.Agresti A. Modelling patterns of agreement and disagreement [published online ahead of print 1992/01/01]. Stat Methods Med Res 1992;1(2):201–218. [DOI] [PubMed] [Google Scholar]
  • 63.Obuchowski NA, Subhas N, Schoenhagen P. Testing for Interchangeability of Imaging Tests. Academic Radiology 2014;21(11):1483–1489. [DOI] [PubMed] [Google Scholar]
  • 64.Hadgu A, Dendukuri N, Wang L. Evaluation of screening tests for detecting Chlamydia trachomatis: bias associated with the patient-infected-status algorithm [published online ahead of print 2011/12/14]. Epidemiology 2012;23(1):72–82. [DOI] [PubMed] [Google Scholar]
  • 65.USFDA. Guidance for industry: Statistical Approaches to Establishing Bioequivalence In. doi: https://www.fda.gov/media/70958/download.: US Food and Drug Administration; 2001. [Google Scholar]
  • 66.Barnhart HX, Kosinski AS, Haber MJ. Assessing individual agreement [published online ahead of print 2007/07/07]. J Biopharm Stat 2007;17(4):697–719. [DOI] [PubMed] [Google Scholar]
  • 67.Obuchowski NA. Can electronic medical images replace hard-copy film? Defining and testing the equivalence of diagnostic tests [published online ahead of print 2001/09/25]. Stat Med 2001;20(19):2845–2863. [DOI] [PubMed] [Google Scholar]
  • 68.World Health Organization (WHO). Cardiovascular diseases (CVDs) Fact Sheet https://www.who.int/en/news-room/fact-sheets/detail/cardiovascular-diseases-(cvds). Accessed 23 April 2020.
  • 69.Lyngbakken MN, Myhre PL, Rosjo H, Omland T. Novel biomarkers of cardiovascular disease: Applications in clinical practice [published online ahead of print 2018/11/21]. Crit Rev Clin Lab Sci 2019;56(1):33–60. [DOI] [PubMed] [Google Scholar]
  • 70.Hafiane A. Vulnerable Plaque, Characteristics, Detection, and Potential Therapies [published online ahead of print 2019/07/31]. J Cardiovasc Dev Dis 2019;6(3). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 71.Stary HC. Natural history and histological classification of atherosclerotic lesions: an update [published online ahead of print 2000/05/16]. Arteriosclerosis, thrombosis, and vascular biology 2000;20(5):1177–1178. [DOI] [PubMed] [Google Scholar]
  • 72.Virmani R, Kolodgie FD, Burke AP, Farb A, Schwartz SM. Lessons from sudden coronary death: a comprehensive morphological classification scheme for atherosclerotic lesions [published online ahead of print 2000/05/16]. Arteriosclerosis, thrombosis, and vascular biology 2000;20(5):1262–1275. [DOI] [PubMed] [Google Scholar]
  • 73.Virmani R, Burke AP, Farb A, Kolodgie FD. Pathology of the Vulnerable Plaque. JACC 2006;47(8):C13–18. [DOI] [PubMed] [Google Scholar]
  • 74.Neglia D, Rovai D, Caselli C, et al. Detection of significant coronary artery disease by noninvasive anatomical and functional imaging. Circ Cardiovasc Imaging 2015;8(3). [DOI] [PubMed] [Google Scholar]
  • 75.Williams MC, Moss AJ, Dweck M, et al. Coronary Artery Plaque Characteristics Associated With Adverse Outcomes in the SCOT-HEART Study. J Am Coll Cardiol 2019;73(3):291–301. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 76.ISCHEMIA: Invasive Strategy No Better Than Meds for CV Events @tctmd; 2020. https://www.tctmd.com/news/ischemia-invasive-strategy-no-better-meds-cv-events.
  • 77.Buckler AJ, Karlöf E, Lengquist M, et al. Virtual Transcriptomics: Noninvasive Phenotyping of Atherosclerosis by Decoding Plaque Biology From Computed Tomography Angiography Imaging [published online ahead of print 2021/03/12]. Arteriosclerosis, thrombosis, and vascular biology 2021. doi: 10.1161/atvbaha.121.315969:Atvbaha121315969. [DOI] [PMC free article] [PubMed]
  • 78.Sheahan M, Ma X, Paik D, et al. Atherosclerotic Plaque Tissue: Noninvasive Quantitative Assessment of Characteristics with Software-aided Measurements from Conventional CT Angiography [published online ahead of print August 31, 2017]. Radiology 2018;286(2):622–631. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 79.QIBA Computed Tomography Angiography Biomarkers Committee. QIBA Profile: Atherosclerosis Biomarkers by Computed Tomography Angiography (CTA)-2020 Profile Stage: Consensus http://qibawiki.rsna.org/images/8/87/QIBA_CTA_Profile_as_of_2020-Mar-10.pdf.
  • 80.Fleiss JL, Levin B, Paik MC. Statistical methods for rates and proportions Third Edition ed: Wiley; 2013. [Google Scholar]
  • 81.Bankier AA, Levine D, Halpern EF, Kressel HY. Consensus interpretation in imaging research: is there a better way? [published online ahead of print 2010/09/21]. Radiology 2010;257(1):14–17. [DOI] [PubMed] [Google Scholar]
  • 82.EP29-A. CAGCd. Expression of Measurement Uncertainty in Laboratory Medicine Wayne, PA: Clinical and Laboratory Standards Institute; 2012. [Google Scholar]
  • 83.Rudin C Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead. Nature Machine Intelligence 2019;1(5):206–215. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 84.Liao CC, Chen YF, Xiao F. Brain Midline Shift Measurement and Its Automation: A Review of Techniques and Algorithms [published online ahead of print 2018/06/01]. Int J Biomed Imaging 2018;2018:4303161. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 85.Mello VSR. Cervical Intraepithelial Neoplasia.. StatPearls [Internet] Treasure Island (FL): StatPearls Publishing; 2022. \; 2022. [Google Scholar]
  • 86.Jackson CH SL, Thompson SG, Duffy SW, Couto E. Multistate Markov Models for Disease Progression with Classification Error. J Royal Statistical Society Series D (The Statistician), 2003;52 193–209. [Google Scholar]
  • 87.Bossuyt PM, Reitsma JB, Linnet K, Moons KG. Beyond diagnostic accuracy: the clinical utility of diagnostic tests [published online ahead of print 2012/06/26]. Clin Chem 2012;58(12):1636–1643. [DOI] [PubMed] [Google Scholar]
  • 88.Fryback DG, Thornbury JR. The efficacy of diagnostic imaging [published online ahead of print 1991/04/01]. Med Decis Making 1991;11(2):88–94. [DOI] [PubMed] [Google Scholar]
  • 89.Bossuyt PMM, Lijmer JG, Mol BWJ. Randomised comparisons of medical tests: sometimes invalid, not always efficient. The Lancet 2000;356(9244):1844–1847. [DOI] [PubMed] [Google Scholar]
  • 90.Simon R Clinical trial designs for evaluating the medical utility of prognostic and predictive biomarkers in oncology [published online ahead of print 2010/04/13]. Per Med 2010;7(1):33–47. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 91.Marsh TL, Janes H, Pepe MS. Statistical inference for net benefit measures in biomarker validation studies. Biometrics 2019;76(3):843–852. [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

1

RESOURCES