Skip to main content
NIHPA Author Manuscripts logoLink to NIHPA Author Manuscripts
. Author manuscript; available in PMC: 2020 Dec 7.
Published in final edited form as: J Proteome Res. 2020 Apr 7;19(7):2794–2806. doi: 10.1021/acs.jproteome.0c00118

Identification of Putative Early Atherosclerosis Biomarkers by Unsupervised Deconvolution of Heterogeneous Vascular Proteomes

Sarah J Parker 1, Lulu Chen 2, Weston Spivia 1, Georgia Saylor 3, Chunhong Mao 4, Vidya Venkatraman 1, Ronald J Holewinski 1, Mitra Mastali 1, Rakhi Pandey 1, Grace Athas 5, Guoqiang Yu 2, Qin Fu 1, Dana Troxlair 5, Richard Vander Heide 5, David Herrington 3,*, Jennifer E Van Eyk 1,*, Yue Wang 2,*
PMCID: PMC7720636  NIHMSID: NIHMS1642712  PMID: 32202800

Abstract

Coronary artery disease remains a leading cause of death in industrialized nations, and early detection of disease is a critical intervention target in order to effectively treat patients and manage risk. Proteomic analysis of mixed tissue homogenates may obscure subtle protein changes that occur uniquely in underlying tissue subtypes. The unsupervised ‘convex analysis of mixtures’ (CAM) tool has previously been shown to effectively segregate cellular subtypes from mixed expression data. In this study, we hypothesized that CAM would identify proteomic information specifically informative to early atherosclerosis lesion involvement that could lead to potential markers of early disease detection. We quantified the proteome of 99 paired Abdominal Aorta (AA) and Left Anterior Descending Coronary Artery (LAD) specimens (N=198 specimens total) acquired during autopsy of young adults free of diagnosed cardiac disease. The CAM tool was then used to segregate protein subsets uniquely associated with different underlying tissue types, yielding markers of normal and fibrous plaque (FP) tissues in LAD and AA (N=62 lesions markers). CAM-derived FP marker expression was validated against pathologist estimated luminal surface involvement of FP, as well as in an orthogonal cohort of ‘pure’ fibrous plaque, fatty streak, and normal vascular specimens. A targeted mass spectrometry (MS) assay quantified 39 of 62 CAM-FP markers in plasma from women with angiographically verified coronary artery disease (CAD, N=46) or free from apparent CAD (control, N=40). Elastic net variable selection with logistic regression reduced this list to 10 proteins capable of classifying CAD status in this cohort with <6% misclassification error, and a mean area under the receiver operating characteristic curve of 0.992 (Confidence Interval 0.968-0.998) after cross validation. The proteomics-CAM workflow identified lesion-specific molecular biomarker candidates by distilling the most representative molecules from heterogenous tissue types.

Keywords: Atherosclerosis, proteomics, DIA-MS, MRM-MS, convex analysis of mixtures

Graphical Abstract

graphic file with name nihms-1642712-f0007.jpg

Introduction

Atherosclerosis is an initially indolent vascular condition of progressive remodeling in the arterial intima and media. The process involves the early accumulation of fatty acids and cholesterols, defined as fatty streaks, within the resident and infiltrating (e.g., macrophages) cells of a vascular region. Subsequent inflammation and immune-system involvement drives development of later stage fibrous plaques which protrude into the vessel lumen 1. Ulceration or rupture of the accumulating lesion and subsequent acute thrombus formation account for the clinical manifestations of atherosclerosis such as myocardial ischemia and infarction, stroke, and peripheral limb disease. Among these, coronary artery disease (CAD) remains the leading cause of death in the United States. Despite its widespread prevalence, appropriate risk-stratification of CAD remains a substantial clinical challenge, with estimates as high as >60% of major adverse cardiac events occurring among individuals deemed low risk based on traditional models2. The effective detection of atherosclerotic status and burden in unselected populations by circulating factors remains limited, with a pressing need for sensitive and specific circulating biomarkers of disease status to more accurately estimate risk of major adverse cardiac events and target high risk individuals for early preventive interventions.

A long-standing yet challenging goal in proteome research is the identification of diagnostic and prognostic protein biomarkers that can improve the standard of care for clinical disease management3. Identifying candidate biomarkers directly from the affected tissue, and ideally at an early stage in disease progression may provide the best opportunity to find proteins that are highly correlated with pathology and that are both sensitive and specific for the presence of lesions4. Proteomic profiling of tissues is often performed by homogenizing tissue samples which may include varying proportions of different cell or tissue types. This produces a mixture proteomic read out in which the informative subtypes and their distinct molecular profiles is lost due to confounding or dilutional effects of other subtypes5. More discriminating alternatives exist, such as separation of tissue types by dissociation and cell sorting or dissection/microdissection of sub-populations, however these methods require a priori knowledge of expected tissue subtypes and the capacity to distinguish between them in a reliable and reproducible way. These techniques are also more labor intensive and expensive. An alternative is to perform computational deconvolution. While most previously reported computational methods for the deconvolution of complex molecular signals require a priori knowledge of expected tissue composition and/or sentinel marker proteins or genes69, the recently introduced Convex Analysis of Mixtures (CAM) method is capable of unsupervised deconvolution of the signals within complex molecular mixtures into tissue-subtype specific molecular expression10. Recent publications have validated the capability of CAM to detect and characterize the heterogeneity within, and identify subtype marker genes of, both real and experimentally produced molecular mixtures from RNA-transcriptome data10. The use of an unsupervised approach would also have the potential to identify novel tissue-markers of pathology without bias.

Recent publications interrogating the proteome of atherosclerosis have mostly focused on late stage disease in the carotid and coronary artery, and few have been conducted in both coronary arteries and abdominal aorta (e.g., muscular artery versus conduit artery)1122. Furthermore, none of the aforementioned studies attempted to computationally segregate protein source based on tissue subtypes present in heterogeneous arterial homogenates. Recently, we analyzed the proteome of human left anterior descending coronary artery (LAD) and abdominal aorta (AA) with varying degrees of lesion involvement as quantified by label-free spectral counting of Data Dependent Acquisition (DDA) Mass Spectrometry (MS) data. Multiple bioinformatic analyses revealed several novel characteristics about the fundamental molecular architecture of the vasculature as lesions develop, and that clear differences in molecular phenotype exist between the LAD and AA. While we compiled a list of potential biomarkers from those analyses and demonstrated their utility in detecting CAD in clinical plasma samples, this application was not the direct focus of that work. In the current study, our goal was to the test the validity of the CAM approach as a stand-alone method for identifying salient protein markers of atherosclerosis across specimens with a continuum of lesion severity. To facilitate rapid transition from discovery into translation, we employed Data Independent Acquisition (DIA) MS, for which a targeted peptide peak group readily exists for each peptide identified from a protein of interest detected in discovery phase, and thus these targets can be quickly vetted for plasma detectability using DIA-MS maps of human plasma samples. Therefore, in this study we demonstrate the utility of applying the CAM tool to DIA-MS arterial proteomic profiles and present a workflow for the identification and preliminary validation of tissue-derived protein biomarkers of atherosclerotic vascular remodeling in a cohort of relatively young subjects (<50 years of age) free from known clinical cardiovascular disease.

METHODS

Overall Study Design.

In order to introduce CAM as a tool for the analysis of DIA-MS proteomic data, and to test the applicability of this approach to putative biomarker development, we completed a unique study employing three distinct MS experiments on unique and orthogonal sample types (Figure 1).

Figure 1.

Figure 1.

Schematic Overview of Study Design. LAD=Left Anterior Descending Coronary Artery. Yellow boxes indicate samples analyzed by Data Independent Acquisition Mass Spectrometry (DIA-MS). The red box indicates samples analyzed by targeted Multiple Reaction Monitoring Mass Spectrometry (MRM-MS). CAD = Coronary Artery Disease.

First, we performed a DIA-MS analysis on human vascular tissues from the coronary artery and aorta in which the amount of lesion (fibrous plaque or fatty streak) varied on a spectrum from 0-100% (based on pathologist estimates of luminal surface involvement, Figure 2). We performed the unsupervised CAM analysis on these tissues-of-heterogenous-disease-status with the net output being a set of proteins identified as ‘Fibrous Plaque’ markers. Second, to validate the accuracy of our inference that these proteins were uniquely derived from the fibrous plaque component of the discovery vascular specimen, we procured an orthogonal validation set from different human subjects of vascular tissue dissected by pathologists so that the specimens were solely composed of zero lesion, pure fatty streak, or pure fibrous plaque. We performed a second DIA-MS experiment on these samples and compared the quantitative and qualitative expression of the putative FP markers between the different pure lesion samples. Thirdly, as a separate effort designed to validate whether the FP marker proteins of interest, identified from tissues, could be (1) detected and (2) biologically discriminative in a test cohort of human plasma, we performed a preliminary MRM-MS experiment characterizing their expression and relative abundance in completely separate samples of plasma from women with or without verified coronary artery disease. The goal was not to produce a fully mature proteomic profile for prediction of coronary disease, but rather to verify that CAM applied to DIA proteomic profiles of mixed tissue samples is indeed capable of identifying subsets of the detected proteome that are sensitive to the presence of coronary atherosclerosis.

Figure 2. Characteristics of the discovery cohort.

Figure 2.

(A) Schematic of vascular specimens used. (B) Table describing the age, sex, and race of the vascular donors in the heterogeneous discovery cohort (N=100). (C) Total composition maps for each specimen of AA and LAD from all 100 donors. Yellow = Fatty streak, Orange = Fibrous Plaque, Grey = Normal.

Subjects and specimens

Vascular tissue for the initial discovery phase was obtained during autopsy from male and female subjects of any race without any prior diagnosis of cardiovascular disease within 48 hours of death. The Medico-Legal Death Investigators obtain signed family informed consent for retrieval of anatomic specimens prior to the autopsy. In total, up to 1 gram segments were collected from two standardized regions of the left anterior descending coronary artery (LAD), the abdominal aorta (AA), and one segment of thoracic aorta. A trained pathologist scored each specimen for surface involvement of fatty streak or fibrous plaque / calcified lesions. Specimens were flash frozen in liquid nitrogen and stored at −80 degrees Celsius until further processing. For the tissue-based validation phase additional, targeted pure specimens of 100% normal (NL, N=3), 100% fatty streak (FS, N=3) or 100% fibrous plaque (FP, N=4) samples were collected, stored and processed in a similar manner as above from separate autopsy patients.

Protein extraction and preparation for mass spectrometry analysis

Unselected (e.g., mixed lesion burden) Arterial Specimens.

Each paired distal aorta and mid LAD tissue was pulverized in liquid nitrogen and homogenized in 8M urea, 2M thiourea, 4% CHAPs and 1% DTT using a dounce homogenizer for 100 strokes. Debris were pelleted by 20 minutes centrifugation at 16,000 rpm. Protein concentration of the supernatant was assessed by CB-X assay kit (G-Biosciences MO, USA) and 15 μg of protein was aliquoted and precipitated using 2-D clean-up kit (GE Healthcare MA, USA) and then reconstitute in 6M urea, 50mM ammonium bicarbonate. The protein was reduced, alkylated and digested with trypsin (1:20) overnight at 37 degrees Celsius. Peptides were desalted using Oasis HLB 96-well plate (Waters MA, USA). To generate a peptide spectral library for subsequent identification and quantification of peptides and proteins, peptides from representative specimen were pooled and separated into 80 basic reverse phase fractions, which were subsequently strategically recombined from distally eluting segments into 24 fractionated peptide samples. These were analyzed by data dependent acquisition mass spectrometry analysis for the assembly of a human vascular peptide assay library.

Pure Pathological Specimen.

Pure specimens were pulverized under liquid nitrogen and solubilized in 8M UREA. Protein was extracted from 100μg of this lysate using high pressure extraction on a Pressure Biosciences 2320 Ext Barocycler, with a protocol of 60 cycles of: 50s at 45 kPSI followed by 10s at atmospheric pressure, performed in micropestle tubes. Protein concentration was determined by BCA assay (Pierce) reduced with DTT, alkylated with IAA and diluted to < 2M Urea concentration in TRIS buffer, with pressure-facilitated tryptic digestion was then performed in the barocycler (45 kPSI 50s, ATM 10s, 90 cycles at 37˚C). Resulting peptides were desalted on waters HLB desalting plates and dried for storage (−80oC) prior to DIA-MS analysis.

MS Acquisition, Peptide Identification, and Quantification Data Analysis Pipeline

MS data were acquired and analyzed as described in previous reports2324 and detailed below:

Mass Spectrometry Acquisition

Data Independent Acquisition.

Four ug of peptides from each unselected analytical specimen, as well as the 10 ‘pure’ specimen, were spiked with exogenous retention time peptide standards (iRT, Biognosys) directly before being loaded on an Eksigent 415 HPLC system operating in microflow mode equipped with a Ekspert nanoLC 400 autosampler. Peptides were first loaded onto a trap column (10×0.3 mm, C18CL, 5 μm, 120Å, Sciex) for 3 minutes at 10 μL/min of solvent A (0.1% FA in water) followed by separation on an analytical column (ChromXP C18CL, 150×0.3 mm, 3μm, 120Å, Sciex) at a flow rate of 5 μL/min using a linear AB gradient of 3-35% solvent B (0.1% FA in ACN) for 60 minutes, 35-85% B for 2 minutes, holding at 85% B for 5 minutes, then re-equilibrating at 3% B for 7 minutes. Mass spectra were collected in data independent acquisition mode, with the instrument looping through an initial MS1 scan of 250 ms ranging from 400-1250 m/z followed by the acquisition of 100 MS2 scans of 30 ms ranging from 100-1800 m/z from peptide ions filtered through mass windows of variable width (isolation window set up provided as a Supplementary Table). Total cycle time was 3.3 seconds, allowing for a minimum of 10 sampled points across a given chromatographic peak for subsequent quantification of peptide MS2 fragments. Source gas 1 was set to 15, gas 2 was set to 20, curtain gas set to 25, source temperature set to 100 °C, and source voltage set to 5500 V.

Peptide Identification and Quantification Data Analysis Pipeline

Targeted Analysis of DIA-MS files.

Peptide peak groups were extracted from an existing library of pooled human vascular lysates described previously11 and found here: http://www.peptideatlas.org/PASS/PASS01066 using the openSWATH workflow2526 implemented in house as described previously24. Briefly, peptide peak groups are extracted from raw DIA-MS data in profile mzML format, and putative identifications are scored according to multiple relevant metrics. Inclusion of decoy peak groups in the peptide assay library allows for modeling of the resulting score distributions and assignment of false discovery rates to target peptide identifications. The TRIC algorithm 27 was then used to align all peptide identifications across the full unselected specimen acquisition sets, which was performed separately within the LAD and the AA specimen cohorts. Fragment level area under the curve data from each file were normalized to the total extracted MS2 signal intensity of that file (e.g., akin to normalization to total protein in a western blot) and normalized fragment intensity data were input into the mapDIA software28 for selection of high quality fragments and aggregation of fragment level data into peptide intensities, and subsequent peptide level intensities into estimates of protein level intensity/abundance.

Application of Convex Analysis of Mixtures across the DIA-MS protein data from unselected specimens.

Protein data were analyzed through the CAM pipeline, as described in 10, 29. Briefly, with the linear latent variable model having been extensively applied to address tissue heterogeneity 7, 30, CAM exploits the strong parallelism between this linear latent variable model and the theory of convex sets to achieve fully unsupervised data deconvolution31. As a fully unsupervised method, the CAM algorithm analyzes the scatter simplex of globally measured protein expression, and geometrically identifies the vertices and their resident molecular markers aggregated from the entire sample cohort. Using these de novo markers, CAM then estimates the proportions of each identified putative tissue subtypes within individual samples. Finally, CAM also estimates the average expression profile over all samples of each protein identified in each putative tissue subtype. The specific steps are as follows: Preprocessing of the data: proteins whose signal intensity (vector norm) is lower than 5% (noise) or higher than 95% (outlier) of the mean value over all proteins are first eliminated as unreliable and likely to skew the subsequent deconvolution; principal component analysis encompassing 10 PCs is then performed on raw measurements, with subsequent aggregation of protein vectors into representative clusters using affinity propagation clustering (APC) to further reduce impact of noise/outlier 32. Then, Minimum Descriptive Length (MDL) parameterization is applied for determining the number of putative tissue subtypes10. The relative proportions of constituent tissues are then estimated by standardized averaging of the expression levels of CAM-identified tissue specific marker proteins. The resulting proportions are then used to deconvolute the mixed expressions into tissue-specific profiles by non-negative least-square regression techniques10. In other words, after preprocessing to reduce outlier interference, the scatter simplex of varying size (number of vertices) is identified, and the MDL parameterization step is run to identify the proper number of sources corresponding to the lowest MDL. The proteins with exclusive or near exclusive expression in just one source are designated as a marker for that source. The CAM pipeline then estimates relative proportions for each of the specified number of sources, and the corresponding averaged expression profile of a given protein is calculated across each of the sources specified. We then used gene ontology analysis and literature search to identify the functional or cell-type specific nature of the marker proteins identified and provide an overall biological classification of a given subtype once it was identified by CAM.

Scheduled Multiple Reaction Monitoring (MRM) Build and Plasma Sample Preparation for the Clinical Plasma Validation Experiment

The targeted peptide assay library entries for the 62 fibrous plaque proteins of interest were queried in Skyline33 against DIA-MS acquisitions of undepleted, control human plasma pool digests produced in our laboratory for quality control monitoring. Strong performing peptides, ideally free of methionine or cysteine residues, were manually selected based on the coelution of fragments, overall intensity, signal to noise, and proximity to expected chromatographic retention time based on iRT prediction. From this list, a targeted method was assembled containing 249 fragments that defined 73 peptides uniquely identifying 42 of the original 62 FP-marker proteins from the tissue discovery analysis. Multiple unscheduled acquisition methods were constructed to determine peptide retention time, and these were input into a final scheduled acquisition method monitoring each fragment within a 2-minute window of its expected elution time.

Plasma specimen from a cohort of older women with angiographically verified coronary artery disease (one or more lesions >30% stenosis) (N=46) obtained after receipt of informed consent from participants in a randomized trial of estrogen replacement therapy as described previously 34 were used to examine FP-marker expression in disease (CAD). Plasma from a group of age-matched women with no known cardiovascular disease (N=40, mean age 60±4 years) obtained through the Cedars-Sinai biobank were used as controls (CONTROL).

To compare abundance of these 42 proteins between women with and without CAD, peptides from undepleted plasma were prepared on a fully automated Biomek NXP Span-8 Laboratory Automation Workstation (Beckman Coulter). Specifically, 5 μL plasma from each sample was diluted in 27.5 μL buffer (0.1M Tris pH 8.3, 4 mM CaCl2) with 5 μL Denaturant (20% N-octyl-glucoside, OGS), 10 μg β-gal in 5 μL and 5 μL Reducing Reagent (50 mM TCEP) and digestion plates were shaken at 1000 RPM at 60 ˚C for 60 minutes. Next, 2.5 μL MMTS (200 mM), followed by an additional 10 minutes shaking at 1000 RPM. Finally, 10 μL trypsin (2.5 μg/μL trypsin in 0.1% FA) was added samples were incubated with shaking at 43˚C for 2 hours. Peptides were separated on a Prominence UFLCXR HPLC system (Shimadzu, Japan) with a Waters Xbridge BEH30 C18 2.1mm x 100mm, 3.5μm column (Waters) flowing at 0.25 mL/min and 36 °C coupled to a QTRAP® 6500. Mobile phase A consisted of 2% ACN, 98% water, and 0.1% formic acid and mobile phase B of 95% ACN, 5% water, and 0.1% formic acid. After loading, the column was equilibrated with 5% B for 5 minutes. Peptides were then eluted over 30 minutes with a linear 5% to 35% gradient of buffer B. The column was washed with 98% B for 10 minutes and then returned to 5% B for 5 minutes before loading the next sample. Scheduled, multiplexed MRM acquisitions for each of the 249 fragments representing 73 peptides from 42 target proteins were completed on each plasma sample, with an additional 10 samples of pooled plasma digest used to verify fragment quantitative quality and acquisition quality. While previously synthesized, heavy SIS peptides were spiked-in for a handful of proteins, these standards were not used in the current analysis. Raw data were imported into Skyline software33 where peaks were manually verified and the area under the intensity x chromatographic elution curve was calculated for each fragment. Fragments with <20% coefficient of variation across the pooled sample acquisitions were used for quantification. Following this filtering step, all fragments for a given protein were averaged to yield a final protein level intensity value. The full MRM experiment, with raw chromatogram files, are publicly available on the Panorama web server (https://panoramaweb.org/CAM_AtheroFPproteomics.url).

Statistical Analysis of CAM Identified Proteins in the Pure Tissue and in the Plasma Validation Samples.

Boxplots plots of protein intensities and the correlation-matrix plot were produced in R using the ggplot2 package35. The pairwise differences between the pure FP, pure FS, and pure normal specimen were calculated using the mapDIA specialized model-based approach, described by its authors as utilizing a Bayesian latent variable model with Markov random field prior. Using this approach, multiple testing is adjusted for with Bayesian False Discovery Rate estimates, or q-values, as described in detail by the mapDIA authors28.

The JMP package within the SAS software suite (SAS, Cary, NC) was used to perform adaptive elastic net logistic regression (α=0.9) to select plasma proteins most predictive of CAD-status within the full patient cohort The mean value of the log-likelihood from a 5-fold cross validation procedure was used to select the tuning parameter for the final model. The confidence limits for the AUC of the final model were derived based on Baysian bootstrap resampling with fractional weights (n=10,000). Figures were created using the R package ggplot2 and Microsoft Excel, with editing of non-scientific content for readability and formatting in Adobe Photoshop.

RESULTS

Data Independent Acquisition MS Analysis of Human Vascular Tissue Specimens

The characteristics of the 100 subjects and their AA and LAD vascular specimens are presented in Figure 2A and B. The data were biased toward higher normal composition (Figure 2C), with higher FS burden detected in the AA relative to the LAD (mean %FS across specimen 20.26 ± 22.8 vs 7.2 ± 16.5 respectively, p[independent, two-tailed ] < 0.00001), whereas fibrous plaque burden was roughly equal between the two vascular regions (10.3 ± 23.8 vs 9.15 ± 21.8 respectively, p[independent, two-tailed]=0.72). One and two samples were lost during processing from the left anterior descending coronary artery (LAD) and abdominal aorta (AA), respectively, so that the final extraction of the DIA-MS data was from N=98 AA and N=99 LAD tissue specimen. In total, 10,809 peptides from 1,744 proteins from AA and 12,371 peptides from 2,055 proteins in LAD samples were quantified from the pooled human vascular library, of which 1274 and 1582 proteins were identified with <50% missingness across all specimen and thereby included in subsequent analyses (Supplementary Tables 1 and 2A, with peptide counts for each protein listed in Supplementary Tables 1 and 2B). Principal Component Analysis demonstrated only marginal grouping of samples with high FP burden, with little clustering among other samples (Supplementary Figure 2A), likely due to the heterogenous degree of lesion burden in the samples.

Analysis of Heterogenous Vascular Specimen using Convex Analysis of Mixtures to identify Tissue Subtypes and their Marker Proteins.

Unsupervised CAM analysis was performed to estimate the presence of likely tissue subtypes and identify the marker proteins that most clearly define them. The CAM analysis indicated two and four distinct expression subtypes in the AA and LAD specimen, respectively (Figure 3A and Supplemental Figure 1). Literature searches on the individual marker proteins identified by CAM led to the biological categorization of each CAM-identified subtype, with one or three types of ‘normal’ arterial tissue subtypes (e.g., medial layer smooth muscle cell proteins) identified in the AA and LAD, respectively, as well as a ‘fibrous plaque’ (FP) type identified in each (Figure 3B and Supplemental Table 3). There was overlap between many of the three LAD ‘normal’ subtypes and the single ‘normal’ subtype in AA, as well as overlap between the LAD and AA ‘FP’ subtype, however none of the markers of FP in either tissue region overlapped with markers designated to ‘normal’ subtypes in the other region (Figure 3C).

Figure 3. CAM analysis results.

Figure 3.

(A) CAM-scatter simplex plots from the analysis of the AA (left panel) and LAD (right panel) proteomes demonstrating the identification of 2 putative subtypes in the AA and 4 putative subtypes in the LAD and specimen. (B) General categorizations assigned to subtypes based on literature search of representative marker proteins (C) Overlap between marker sets from the 2 AA and 4 LAD subtypes as identified by the CAM analysis. Subtypes are listed as row labels in the table beneath the large bar chart. Small horizontal bars to the left of each row label represent the number of proteins identified as ‘markers’ in each subtype. The larger vertical bar graph represents the number of proteins identified as markers in a given single subtype or overlapping subtypes between an LAD and AA, as denoted by the filled circles without (proteins unique to just that subtype) or with (proteins overlapping between two subtypes) lines connecting any two rows.

The abundance of the 62 putative FP-markers identified by CAM across the unselected LAD and AA specimens indicated a clear enrichment of FP-marker abundance in specimens with higher FP burden as estimated by either the CAM-algorithm or an actual human pathologist grading (Figure 4A). The pattern of putative FP marker expression and pathologist estimated lesion burden was not perfect, with some samples demonstrating inconsistency between pathologist estimated lesion burden and CAM estimated FP-burden, especially in the LAD samples where at least 4 specimen were graded at >60% FP by the pathologist but CAM estimated 40% lesion involvement. Importantly, however, overall the distribution of CAM-estimated FP-subtype composition was lowest among samples with no pathologist graded surface FP and the highest among samples with >50% pathologist graded surface FP (Figure 4B).

Figure 4. Association between CAM-FP Marker Protein Expression and Pathologist Estimated Luminal FP Involvement.

Figure 4.

(A) Expression of the FP marker proteins across all 99 tissue specimens, organized by the origin of the marker as unique to the CAM analysis of AA or LAD separately, or found as a common marker of FP in both tissue types. Expression is displayed in the heatmap as the Z-Score, or Standard Deviations away from the mean, for a given protein (columns) in a given sample (rows). Z-scores were calculated separately for the AA and LAD specimen sets. The specimens (rows) are plotted in order of CAM-estimated fibrous plaque involvement (black bars at left of each heatmap). Pathologist estimated surface FP grading is shown using red bars. (B) Box plots demonstrating the distribution of CAM-FP estimates for specimen categorized by the pathologist surface grading as having no FP, up to 49.9% FP, or ≥ 50% FP.

Validation of CAM-Identified Markers in Pure Pathological Specimens

To validate the FP origin of the marker proteins identified by CAM, we performed DIA-MS analysis of specimens isolated from an orthogonal set of aortic tissue samples specifically isolated from regions considered ‘pure’ for FP, fatty streak (FS), or normal (NL) pathology. In Principal component (PC) analysis, the FP specimen could be separated from both FS and NL specimen along PC1, with further segregation between FS and NL on PC4 (Supplementary Figure 2). Qualitative analysis of the expression patterns of 58 of the 62 putative markers detected in this second analysis corroborates the designation of the CAM-FP markers as indicative of FP (Figure 5B). There were 677 proteins with significantly different expression between pure-FP and pure-NL tissue specimen (FDR < 5%), among which 83% of the 58 CAM-marker proteins (P=48) were either statistically significantly higher expression in pure-FP or untestable expression differences due to lack of- or sparse-detection in the normal samples. While most of the marker proteins were also elevated in FS relative to NL, many were to a lesser extent and fewer were statistically significant (Supplementary Table 5). Finally, the functional groupings of the putative FP proteins were also consistent with atherosclerosis related functional ontologies, including lipid transport (e.g., APOB, APOE, APOL1, APOC1 and APOC2), fibrin assembly (e.g., fibrinogen chains A, B, and G), innate immunity (e.g., IGHM, IGHA2, CD14, S100A8), and the complement cascade (e.g., F10, C4BPA and C4BPB) (Figure 5B and Supplemental Table 4).

Figure 5. Validation of CAM identified FP Markers in Pure Specimen.

Figure 5.

(A) Schematic of ‘pure’ vascular specimens used for validation experiments. (B) Expression heatmap of each of the CAM-FP markers within pure specimen of fibrous plaque, fatty streak, or normal vascular tissue sampled from a separate cohort of young human subjects. False Discovery Rate (FDR) values from the mapDIA comparison of FP vs Normal tissues are shown on the left. (C) Map depicting some of the key function ontology groupings identified among the FP marker set. David Gene Ontology analysis identified multiple clusters of atherosclerosis related functional groupings (shown in black nodes). Quantified, CAM-identified FP markers linked to each of these categories are shown as nodes, with average expression of FS and NL groups displayed as relative to the FP expression for that node. Node bar charts are colored-coded by pure NL (blue), FS (turquoise), and FP (orange) samples. Red boxes denote the proteins that were ultimately selected from plasma MRM data for inclusion in a logistic regression model to classify coronary artery disease status in patients versus controls.

Validation of CAM-identified Tissue FP protein detection in human plasma and Ability to Discriminate Clinical CAD from Control Patients

A major translational goal is to use pathologically discriminative proteins originally identified within affected tissues as circulating biomarkers capable of diagnosing the presence and extent of FP lesions in a patient. To test the feasibility of the CAM-approach to informing possible plasma FP markers, we built a scheduled, multiple reaction monitoring experiment.. This resulting FP-multiplex was then analyzed within 46 post-menopausal women with angiographically verified CAD as well as 40 approximately age matched women with no apparent CAD (Supplemental Figure 3A-B). Thirty-nine proteins demonstrated adequate quantitative precision (Supplemental Figure 3C) and were included in the final statistical analysis. High intercorrelation was observed among these 39 proteins (Supplemental Figure 3D). Thus, we used elastic net variable selection within a logistic regression analysis which selected a panel of 10 proteins (APOM, PLMN, A2MG, IGHA2, APOC2, VTNC, CD14, ITIH2, TSP2, URP2, detailed description in Supplemental text) that parsimoniously classified CAD status within this cohort. Interestingly, these 10 proteins represented much of the spectrum of the functional groups identified from the tissue-based analysis (Figure 4C), and as is permitted in elastic net feature selection, this subset of proteins also represents the groups / clusters of highly inter-correlated proteins (Supplemental Figure 2D). As shown in Figure 6A, by examining the individual intensity distributions for each protein between groups, no one of these proteins was highly efficient at categorizing women with CAD alone. Their linear combination, however, produced a highly accurate classifier with a misclassification rate of 5.8% (95% Confidence Interval 0.057-0.058) (Figure 6B). The expression pattern of these 10 proteins among the women for whom the model predicted the highest (N=5) and lowest (N=5) probability of CAD yielded little insight into a consensus pattern of expression that could uniformly indicate CAD (Figure 6C) from any one of these proteins. Alternatively, the collective logistic model produced a mean bootstrap estimated ROC AUC of 0.992 (95% CI 0.968-0.998) for the classification of women as having CAD in this cohort (Figure 6D).

Figure 6. Initial testing and validation of a preliminary FP-marker panel in human plasma.

Figure 6

(A) Boxplots representing the distribution of abundance for the 10 proteins selected by the elastic net function as the most parsimonious set of classifiers for a logistic regression analysis predicting CAD status. Each individual measurement is shown as a filled circle overlaid on the box plot for that group. (B) Boxplots showing the distribution of the model-derived probability of CAD plotted for the true control and true CAD groups. Sample size is N=40 Control and N=46 CAD plasma specimens. (C) Bar graphs demonstrating the individual protein intensity values for each of the model selected proteins as measured for each of 10 different patients, N=5 with the highest model predicted probability of CAD and N=5 with the lowest model predicted probability of CAD. (D) Receiver operating characteristic (ROC) curve plotting the sensitivity and specificity of the CAD-prediction model.

Discussion

We have quantified the proteome of two vascular regions with variable levels of sub-clinical atherosclerosis and used a novel tool, Convex Analysis of Mixtures, to segregate proteins from these heterogeneous tissue homogenates into putative tissue subtypes including one that generated candidates for biomarkers of fibrous plaque lesions. The marker proteins identified for this subtype showed differential abundance across pure samples of fibrous plaque, fatty streak, and non-lesion abdominal aortic tissues, with many exhibiting some increase in pure FS with a further increase in pure FP relative to normal.

We found significant predictive value in a subset of 10 circulating proteins, derived from the CAM tissue-marker derivation, for classifying CAD-status within a preliminary test cohort. Each of these proteins has been shown in previous studies to be related to atherosclerosis and, in some cases, to have potential as a biomarker of CAD, and a detailed summary of these 10 proteins and existing evidence of their role or connection to atherosclerosis and/or CAD is provided in supplemental text. Importantly, no single protein performed as well independently as the collective model for classifying CAD. A clear advantage of the scheduled targeted mass spectrometry experiment is the simultaneous quantification of large numbers of proteins in one sample36. Capturing multiple potentially informative molecules, and then integrating them into predictive models that can be specialized to address sex, ethnic / genetic, comorbidity, and other environmental or individual modifiers may address the challenges of capturing individual risk and practicing precision medicine. Indeed, our own 23 and other recent work 37 has highlights the value of integrative panels, rather than individual biomarkers, for risk stratification and clinical applications in cardiovascular disease38. We recently published a similar panel of proteins identified using more traditional approaches to acquisition and analysis of proteomic data23. The current study detected 13 proteins that overlapped with our previous work, however the majority (N=26) of high quality, plasma-monitored proteins were unique, and the current set resulted in a stronger performance in logistic regression (0.93 previous versus 0.992 current, 95% CI 0.85-0.97 versus 0.968-0.998, respectively). The unique marker proteins identified in this updated analysis relative to our previous work are likely due to differences in analytical approach for prioritizing candidates and also could have been influenced by the method of proteome acquisition (e.g., DIA here versus DDA previously) or quantification (e.g., aggregated fragment HPLC AUCs here versus spectral counting previously). It is likely that the two different approaches generated complementary information, and future studies will examine the combined marker set in an expanded cohort of CAD patients and controls with additional refinement of the MRM experiment into a fully optimized assay. Furthermore, given (1) the advanced nature of CAD in the test cohort, as well as (2) the observation of many markers with more subtle but still elevated abundance in FS tissues samples, it may be that in a more diverse cohort with pre-clinical or less advanced disease a different set of markers than the 10 prioritized in this analysis will be most informative for disease.

We are presenting a workflow whereby the CAM algorithm is used to identify salient proteins considered highly representative of a pathological tissue subtype. The non-categorical nature of the independent variable in this study (%FP involvement) made traditional statistical analyses challenging to apply to the samples from the initial discovery cohort. While we can apply an arbitrary cut off and group ‘disease’ against ‘non-disease’ samples we risk losing valuable information on protein expression in samples with low or early %FP involvement that fall below selected cut off values. The use of CAM harnesses the heterogeneity of lesion involvement across samples. Other approaches, such as generalized linear modeling, can also apply to this experimental design and have been compared to CAM in our previous work on similar data23, showing that CAM both corroborates but also complements more traditional methods. Additionally, even when grading is the most clear (e.g., pure specimen with 100% lesion involvement), statistical significance in a case-control comparison does not necessarily translate to good ‘marker protein’ potential, and there are often more differentially expressed proteins than are pragmatic for development as marker candidates. For example, the pairwise comparison of pure NL and FP samples in this study identified 677 differentially expression proteins between FP and NL, 188 of which were upregulated more than 1.5 fold, a large number of candidates to sift through and prioritize. Furthermore, simple differential expression does not clarify which proteins have uniquely high expression in just one subtype versus all other subtypes. This may be an important distinction for mechanistic understanding and possibly translation to biomarker development. Deconvolution tools such as CAM can complement more complex designs such as single-cell profiling in future studies into the importance of cell-type or tissue-type of origin. Thus, our data demonstrate that CAM and possibly other deconvolution tools can provide unique insight and contribute to the analysis of heterogenous tissue homogenates in proteomic experiments.

Interestingly, while most samples with high pathologist lesion estimates also had high CAM-estimated FP composition, the CAM estimates of FP grading and the subsequent expression of FP-marker proteins did not always match well with the pathologist grading of FP-lesion burden. Importantly, expression of the CAM-FP marker proteins in pure-specimen with 100% lesion clearly supports their heavily biased distribution in FP lesion tissue. Nevertheless, discrepancy between CAM-estimates and pathologist estimates may indicate that there are some lesions with different molecular phenotypes not captured by the current application of CAM to these specific samples, or that uncertainties both in the visual method or the unsupervised molecular-expression based CAM method for estimating lesion burden will result in discrepancies and some degree of error. Additional studies with even larger sample sizes to increase the diversity of lesion types and involvement will help to answer these questions.

There are several other unsupervised approaches to data deconvolution that may result in complementary or corroborating information, and while CAM has been shown previously to outperform many of these in terms of accuracy of tissue subtype estimation31, it was not our goal in this study to review and compare all possible methods nor to declare CAM the best choice for all applications. Here, we show that with CAM we identified a tissue-derived signature of proteins representative of fibrous plaques present in the vasculature of young, asymptomatic individuals. For the first time with CAM applied to proteomic data, we also validate these CAM-identified putative marker proteins using an orthogonal set of ‘pure’ lesion samples. Our data demonstrate that CAM or other unsupervised deconvolution approaches may be a valuable addition to statistical processing pipelines in biomarker candidate discovery workflows for various disease conditions.

Finally, it is important to note here that this report is not intended to put forth a translation-ready biomarker panel. The plasma work presented herein is meant as preliminary evidence that some of these CAM-identified tissue markers of disease are detectable in the circulation and appear to segregate with the presence of disease, at least in this small-scale preliminary cohort. In the analysis of the plasma expression data, the use of Elastic Net variable selection and cross-validation were both chosen as a means to maximize potential generalizability while working within the constraints of a small, preliminary cohort. It is likely that the predictive power of this model is highly over-estimated and replication in separate, larger and appropriately powered validation cohorts is needed before conclusions regarding true biomarker performance of these preliminary candidates can be drawn. While the sample size of our discovery cohort is one of the largest published to date, it still represents a relatively small sampling of the total population and is thus unlikely to contain the full range of biologically informative molecular information that a larger and more diverse cohort could include. Thus, with extensive additional development these early phase discoveries, combined with those of a growing number of atherosclerosis biomarker studies, may ultimately yield translationally powerful tools for atherosclerosis screening and management.

Supplementary Material

STable3

Supplementary Tables 3A & B – CAM-identified subtype marker protein lists

STable4

Supplementary Tables 4A&B – Gene Ontology and Functional analysis of CAM FP-marker proteins

STable5

Supplementary Tables 5A&B – Protein quantification data and mapDIA pairwise statistical comparison output from pure lesion and pure normal validation samples

STable1

Supplementary Tables 1A & B – Discovery heterogenous lesion containing Abdominal Aorta samples protein quantification results

STable2

Supplementary Tables 2A & B – Discovery heterogenous lesion containing Left Anterior Descending Coronary Artery samples protein quantification results

Supplementary Text and Figures

Supplementary Text and Figures – Functional description and references supporting existing evidence linking the 10 Elastic Net selected plasma proteins to CAD. Supplementary Figures and captions.

Panorama Online Data Repository - DDA library used to search both Discovery and Pure Lesion Validation samples, Chromatograms and raw data for Discovery (LAD and AA) as well as Pure sample DIA-MS (based on openSWATH determined peak integration, scoring, and pyprophet/TRIC FDR modeling and selection)

Acknowledgements.

This work was funded by the Genomics and Proteomic Architecture of Atherosclerosis grant, 5R01HL111362-06

References

  • 1.Stary HC; Chandler AB; Glagov S; Guyton JR; Insull W Jr.; Rosenfeld ME; Schaffer SA; Schwartz CJ; Wagner WD; Wissler RW, A definition of initial, fatty streak, and intermediate lesions of atherosclerosis. A report from the Committee on Vascular Lesions of the Council on Arteriosclerosis, American Heart Association. Circulation 1994, 89 (5), 2462–78. [DOI] [PubMed] [Google Scholar]
  • 2.Myerburg RJ; Junttila MJ, Sudden cardiac death caused by coronary heart disease. Circulation 2012, 125 (8), 1043–52. [DOI] [PubMed] [Google Scholar]
  • 3.Ioannidis JPA; Bossuyt PMM, Waste, Leaks, and Failures in the Biomarker Pipeline. Clinical chemistry 2017, 63 (5), 963–972. [DOI] [PubMed] [Google Scholar]
  • 4.Fleg JL; Stone GW; Fayad ZA; Granada JF; Hatsukami TS; Kolodgie FD; Ohayon J; Pettigrew R; Sabatine MS; Tearney GJ; Waxman S; Domanski MJ; Srinivas PR; Narula J, Detection of high-risk atherosclerotic plaque: report of the NHLBI Working Group on current status and future directions. JACC. Cardiovascular imaging 2012, 5 (9), 941–55. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Parker SJ; Raedschelders K; Van Eyk JE, Emerging proteomic technologies for elucidating context-dependent cellular signaling events: A big challenge of tiny proportions. Proteomics 2015, 15 (9), 1486–502. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Newman AM; Liu CL; Green MR; Gentles AJ; Feng W; Xu Y; Hoang CD; Diehn M; Alizadeh AA, Robust enumeration of cell subsets from tissue expression profiles. Nature methods 2015, 12 (5), 453–7. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.Kuhn A; Thu D; Waldvogel HJ; Faull RL; Luthi-Carter R, Population-specific expression analysis (PSEA) reveals molecular changes in diseased brain. Nature methods 2011, 8 (11), 945–7. [DOI] [PubMed] [Google Scholar]
  • 8.Shen-Orr SS; Tibshirani R; Khatri P; Bodian DL; Staedtler F; Perry NM; Hastie T; Sarwal MM; Davis MM; Butte AJ, Cell type-specific gene expression differences in complex tissues. Nature methods 2010, 7 (4), 287–9. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.Moffitt RA; Marayati R; Flate EL; Volmar KE; Loeza SG; Hoadley KA; Rashid NU; Williams LA; Eaton SC; Chung AH; Smyla JK; Anderson JM; Kim HJ; Bentrem DJ; Talamonti MS; Iacobuzio-Donahue CA; Hollingsworth MA; Yeh JJ, Virtual microdissection identifies distinct tumor- and stroma-specific subtypes of pancreatic ductal adenocarcinoma. Nature genetics 2015, 47 (10), 1168–78. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.Wang N; Hoffman EP; Chen L; Chen L; Zhang Z; Liu C; Yu G; Herrington DM; Clarke R; Wang Y, Mathematical modelling of transcriptional heterogeneity identifies novel markers and subpopulations in complex tissues. Scientific reports 2016, 6, 18909. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11.Lee R; Fischer R; Charles PD; Adlam D; Valli A; Di Gleria K; Kharbanda RK; Choudhury RP; Antoniades C; Kessler BM; Channon KM, A novel workflow combining plaque imaging, plaque and plasma proteomics identifies biomarkers of human coronary atherosclerotic plaque disruption. Clin Proteomics 2017, 14, 22. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Ward LJ; Olausson P; Li W; Yuan XM, Proteomics and multivariate modelling reveal sexspecific alterations in distinct regions of human carotid atheroma. Biol Sex Differ 2018, 9 (1), 54. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.Fasehee H; Fakhraee M; Davoudi S; Vali H; Faghihi S, Cancer biomarkers in atherosclerotic plaque: Evidenced from structural and proteomic analyses. Biochem Biophys Res Commun 2019, 509 (3), 687–693. [DOI] [PubMed] [Google Scholar]
  • 14.Langley SR; Willeit K; Didangelos A; Matic LP; Skroblin P; Barallobre-Barreiro J; Lengquist M; Rungger G; Kapustin A; Kedenko L; Molenaar C; Lu R; Barwari T; Suna G; Yin X; Iglseder B; Paulweber B; Willeit P; Shalhoub J; Pasterkamp G; Davies AH; Monaco C; Hedin U; Shanahan CM; Willeit J; Kiechl S; Mayr M, Extracellular matrix proteomics identifies molecular signature of symptomatic carotid plaques. J Clin Invest 2017, 127 (4), 1546–1560. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Liang W; Ward LJ; Karlsson H; Ljunggren SA; Li W; Lindahl M; Yuan XM, Distinctive proteomic profiles among different regions of human carotid plaques in men and women. Scientific reports 2016, 6, 26231. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Hao P; Ren Y; Pasterkamp G; Moll FL; de Kleijn DP; Sze SK, Deep proteomic profiling of human carotid atherosclerotic plaques using multidimensional LC-MS/MS. Proteomics Clin Appl 2014, 8 (7–8), 631–5. [DOI] [PubMed] [Google Scholar]
  • 17.Bagnato C; Thumar J; Mayya V; Hwang SI; Zebroski H; Claffey KP; Haudenschild C; Eng JK; Lundgren DH; Han DK, Proteomics analysis of human coronary atherosclerotic plaque: a feasibility study of direct tissue proteomics by liquid chromatography and tandem mass spectrometry. Mol Cell Proteomics 2007, 6 (6), 1088–102. [DOI] [PubMed] [Google Scholar]
  • 18.de la Cuesta F; Alvarez-Llamas G; Maroto AS; Donado A; Zubiri I; Posada M; Padial LR; Pinto AG; Barderas MG; Vivanco F, A proteomic focus on the alterations occurring at the human atherosclerotic coronary intima.Mol Cell Proteomics 2011, 10 (4), M110 003517. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.Olson FJ; Sihlbom C; Davidsson P; Hulthe J; Fagerberg B; Bergstrom G, Consistent differences in protein distribution along the longitudinal axis in symptomatic carotid atherosclerotic plaques. Biochem Biophys Res Commun 2010, 401 (4), 574–80. [DOI] [PubMed] [Google Scholar]
  • 20.Porcelli B; Ciari I; Felici C; Pagani R; Banfi C; Brioschi M; Giubbolini M; de Donato G; Setacci C; Terzuoli L, Proteomic analysis of atherosclerotic plaque. Biomed Pharmacother 2010, 64(5), 369–72. [DOI] [PubMed] [Google Scholar]
  • 21.Lepedda AJ; Cigliano A; Cherchi GM; Spirito R; Maggioni M; Carta F; Turrini F; Edelstein C; Scanu AM; Formato M, A proteomic approach to differentiate histologically classified stable and unstable plaques from human carotid arteries. Atherosclerosis 2009, 203 (1), 112–8. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Viiri LE; Full LE; Navin TJ; Begum S; Didangelos A; Astola N; Berge RK; Seppala I; Shalhoub J; Franklin IJ; Perretti M; Lehtimaki T; Davies AH; Wait R; Monaco C, Smooth muscle cells in human atherosclerosis: proteomic profiling reveals differences in expression of Annexin A1 and mitochondrial proteins in carotid disease. J Mol Cell Cardiol 2013, 54, 65–72. [DOI] [PubMed] [Google Scholar]
  • 23.Herrington DM; Mao C; Parker SJ; Fu Z; Yu G; Chen L; Venkatraman V; Fu Y; Wang Y; Howard TD; Jun G; Zhao CF; Liu Y; Saylor G; Spivia WR; Athas GB; Troxclair D; Hixson JE; Vander Heide RS; Wang Y; Van Eyk JE, Proteomic Architecture of Human Coronary and Aortic Atherosclerosis. Circulation 2018, 137 (25), 2741–2756. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24.Parker SJ; Venkatraman V; Van Eyk JE, Effect of peptide assay library size and composition in targeted data-independent acquisition-MS analyses. Proteomics 2016, 16 (15–16), 2221–37. [DOI] [PubMed] [Google Scholar]
  • 25.Rost HL; Rosenberger G; Navarro P; Gillet L; Miladinovic SM; Schubert OT; Wolski W; Collins BC; Malmstrom J; Malmstrom L; Aebersold R, OpenSWATH enables automated, targeted analysis of data-independent acquisition MS data. Nature biotechnology 2014, 32 (3), 219–23. [DOI] [PubMed] [Google Scholar]
  • 26.Rost HL; Aebersold R; Schubert OT, Automated SWATH Data Analysis Using Targeted Extraction of Ion Chromatograms. Methods in molecular biology 2017, 1550, 289–307. [DOI] [PubMed] [Google Scholar]
  • 27.Rost HL; Liu Y; D’Agostino G; Zanella M; Navarro P; Rosenberger G; Collins BC; Gillet L; Testa G; Malmstrom L; Aebersold R, TRIC: an automated alignment strategy for reproducible protein quantification in targeted proteomics. Nature methods 2016, 13 (9), 777–83. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28.Teo G; Kim S; Tsou CC; Collins B; Gingras AC; Nesvizhskii AI; Choi H, mapDIA: Preprocessing and statistical analysis of quantitative proteomics data from data independent acquisition mass spectrometry. J Proteomics 2015, 129, 108–120. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29.Wang N; Gong T; Clarke R; Chen L; Shih Ie M; Zhang Z; Levine DA; Xuan J; Wang Y, UNDO: a Bioconductor R package for unsupervised deconvolution of mixed gene expressions in tumor samples. Bioinformatics 2015, 31 (1), 137–9. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30.Guintivano J; Aryee MJ; Kaminsky ZA, A cell epigenotype specific model for the correction of brain cellular heterogeneity bias and its application to age, brain region and major depression. Epigenetics 2013, 8 (3), 290–302. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31.Zhu Y; Wang N; Miller DJ; Wang Y, Convex Analysis of Mixtures for Separating Non-negative Well-grounded Sources. Scientific reports 2016, 6, 38350. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32.Frey BJ; Dueck D, Clustering by passing messages between data points. Science 2007, 315 (5814), 972–6. [DOI] [PubMed] [Google Scholar]
  • 33.MacLean B; Tomazela DM; Shulman N; Chambers M; Finney GL; Frewen B; Kern R; Tabb DL; Liebler DC; MacCoss MJ, Skyline: an open source document editor for creating and analyzing targeted proteomics experiments. Bioinformatics 2010, 26 (7), 966–8. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34.Herrington DM; Reboussin DM; Brosnihan KB; Sharp PC; Shumaker SA; Snyder TE; Furberg CD; Kowalchuk GJ; Stuckey TD; Rogers WJ; Givens DH; Waters D, Effects of estrogen replacement on the progression of coronary-artery atherosclerosis. The New England journal of medicine 2000, 343 (8), 522–9. [DOI] [PubMed] [Google Scholar]
  • 35.Wickam H, ggplot2: Elegant Graphics for Data Analysis. Springer-Verlag: New York, 2009. [Google Scholar]
  • 36.Percy AJ; Chambers AG; Yang J; Hardie DB; Borchers CH, Advances in multiplexed MRM-based protein biomarker quantitation toward clinical utility. Biochim Biophys Acta 2014, 1844 (5), 917–26. [DOI] [PubMed] [Google Scholar]
  • 37.Anwar MA; Dai DL; Wilson-McManus J; Smith D; Francis GA; Borchers CH; McManus BM; Hill JS; Cohen Freue GV, Multiplexed LC-ESI-MRM-MS-based Assay for Identification of Coronary Artery Disease Biomarkers in Human Plasma. Proteomics Clin Appl 2019, e1700111. [DOI] [PubMed] [Google Scholar]
  • 38.Robin X; Turck N; Hainard A; Lisacek F; Sanchez JC; Muller M, Bioinformatics for protein biomarker panel classification: what is needed to bring biomarker panels into in vitro diagnostics? Expert Rev Proteomics 2009, 6 (6), 675–89. [DOI] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

STable3

Supplementary Tables 3A & B – CAM-identified subtype marker protein lists

STable4

Supplementary Tables 4A&B – Gene Ontology and Functional analysis of CAM FP-marker proteins

STable5

Supplementary Tables 5A&B – Protein quantification data and mapDIA pairwise statistical comparison output from pure lesion and pure normal validation samples

STable1

Supplementary Tables 1A & B – Discovery heterogenous lesion containing Abdominal Aorta samples protein quantification results

STable2

Supplementary Tables 2A & B – Discovery heterogenous lesion containing Left Anterior Descending Coronary Artery samples protein quantification results

Supplementary Text and Figures

Supplementary Text and Figures – Functional description and references supporting existing evidence linking the 10 Elastic Net selected plasma proteins to CAD. Supplementary Figures and captions.

Panorama Online Data Repository - DDA library used to search both Discovery and Pure Lesion Validation samples, Chromatograms and raw data for Discovery (LAD and AA) as well as Pure sample DIA-MS (based on openSWATH determined peak integration, scoring, and pyprophet/TRIC FDR modeling and selection)

RESOURCES