Skip to main content
The Breast : Official Journal of the European Society of Mastology logoLink to The Breast : Official Journal of the European Society of Mastology
. 2026 Aug 18;89:104896. doi: 10.1016/j.breast.2026.104896

Revisiting HER2-negative breast cancer in the era of HER2-low and HER2-ultralow expression: Assessment of interobserver variability

Xiao Huang a,1, Sarah Anderson a, Valeria Dal Zotto a, Shuko Harada a,d, Andrea G Kahn a, Meiling Zhou c, Kui Zhang b, Kanako Okamoto a, Shi Wei a,d,⁎
PMCID: PMC13640464  PMID: 42664757

Abstract

Background

HER2-low and HER2-ultralow expressions have recently emerged as actionable targets in breast cancer (BC). However, data remain limited regarding the consistency of immunohistochemistry (IHC) scoring when both categories are considered. This study aimed to evaluate interobserver variability in assessing HER2-negative BC, with particular focus on the reproducibility of HER2-low and -ultralow classifications.

Methods

A total of 389 consecutive HER2-negative BC cases were re-evaluated by six pathologists with varying experience (1-20 years). The panel included dedicated breast pathologists, general pathologists routinely covering breast services, and a surgical pathology fellow in training. HER2 expression was categorized as HER2-null (no staining), -ultralow, 1+, or 2+. Interobserver agreement was analyzed using Fleiss’ kappa statistics and intraclass correlation coefficients (ICCs).

Results

Three scoring schemes were compared: a 4-tier scheme (null, ultralow, 1+, 2+), a 3-tier scheme (combining 1+ and 2+), and a 2-tier scheme (combining ultralow and low). The overall Fleiss’ kappa values were 0.67, 0.68, and 0.77 for the 4-, 3-, and 2-tier schemes, respectively, indicating substantial agreement. Ten rater pairs demonstrated substantial agreement (0.61–0.80), and five showed moderate agreement (0.41–0.60), all involving one rater who consistently assigned lower scores. ICCs ranged from 0.75 to 0.90, reflecting good reliability. Unanimous agreement among all raters was achieved in 53.0%, 61.2%, and 85.7% of cases with the 3 tier schemes, respectively.

Conclusion

Despite the nuanced distinction between HER2-low and HER2-ultralow categories, IHC remains a reliable and reproducible method for evaluating HER2 expression in BC, even in the era of expanded HER2-targeted therapies, although reproducibility is lower for borderline HER2-ultralow cases.

Keywords: Breast cancer, HER2, Interobserver variability, HER2-low, HER2-ultralow

Graphical abstract

HER2-low and HER2-ultralow expressions have recently emerged as actionable targets in breast cancer (BC). However, data remain limited regarding the consistency of immunohistochemistry scoring when both categories are considered. In this study, a total of 389 HER2-negative BC cases were evaluated by six pathologists with varying levels of experience (1–20 years), with particular focus on the reproducibility of HER2-low and HER2-ultralow classifications. The panel included dedicated breast pathologists, general pathologists who routinely interpret breast specimens, and a surgical pathology fellow in training. HER2 expression was categorized as HER2-null (no staining), HER2-ultralow, 1+, or 2+. Three scoring schemes were compared: a 4-tier scheme (null, ultralow, 1+, 2+), a 3-tier scheme (combining 1+ and 2+), and a 2-tier scheme (combining ultralow and low). The overall Fleiss' kappa values were 0.67, 0.68, and 0.77 for the 4-, 3-, and 2-tier schemes, respectively, indicating substantial agreement. Ten rater pairs demonstrated substantial agreement (0.61–0.80), whereas five showed moderate agreement (0.41–0.60), all involving one rater who consistently assigned lower scores. Intraclass correlation coefficients ranged from 0.75 to 0.90, reflecting good reliability. Unanimous agreement among all raters was achieved in 53.0%, 61.2%, and 85.7% of cases using the 4-, 3-, and 2-tier schemes, respectively. Despite the nuanced distinction between HER2-low and HER2-ultralow categories, our findings demonstrate that IHC remains a reliable and reproducible method for evaluating HER2 expression in BC, even in the era of expanded HER2-targeted therapies, although reproducibility is lower for borderline HER2-ultralow cases.

graphic file with name ga1.webp

Highlights

  • •

    The findings reflect real-world practice where the raters include a diverse group of pathologists.

  • •

    Individual raters' pre-existing interpretive thresholds likely contributed to the major discrepancies.

  • •

    An internal consensus for final HER2 score assignment in challenging cases plays a key role in the overall superior interobserver agreement.

  • •

    Lower interobserver consistency is most often observed in cases bordering HER2-null and -ultralow.

1. Introduction

The human epidermal growth factor receptor 2 (HER2), a member of the epidermal growth factor receptor family, is one of the most extensively studied biomarkers in breast cancer. As a receptor tyrosine kinase, HER2 plays a pivotal role in mammary carcinogenesis. While it is expressed at low levels in normal breast epithelium, HER2 overexpression or gene amplification occurs in approximately 15–20% of breast cancers and is historically linked to more aggressive biological behavior [1]. However, the advent of HER2-targeted therapies has significantly transformed the treatment landscape for HER2-positive breast cancer, improving both therapeutic outcomes and prognosis [2].

Over the past two decades, the assessment of HER2 status as a predictive marker for HER2-targeted therapies has primarily focused on distinguishing HER2-positive cancers from HER2-negative tumors. This testing paradigm relies on validated immunohistochemistry (IHC) and in situ hybridization (ISH) assays to evaluate HER2 protein expression and gene amplification, respectively. The guidelines for HER2 testing, established by the American Society of Clinical Oncology (ASCO) and the College of American Pathologists (CAP), have continuously evolved to address changing clinical demands [[3], [4], [5], [6]]. Notably, interobserver agreement for HER2-positive tumors has generally been found to be substantial [7].

The spectrum of HER2 expression extends beyond just positive and negative classifications. HER2-low breast cancer, characterized by an IHC score of 1+ or 2+/non-amplified, is reported to account for 45-55% of all breast cancers [8]. Furthermore, HER2-ultralow breast cancer refers to an IHC score of 0, with some incomplete or faint membranous expression observed in up to 10% of cells [9]. Historically, both HER2-low and HER2-ultralow breast cancers have been classified as HER2-negative, as traditional HER2-targeted therapies are not effective in these cases [10].

Trastuzumab deruxtecan (T-DXd), an antibody-drug conjugate, was recently approved by the U.S. Food and Drug Administration for the treatment of patients with HER2-low and HER2-ultralow metastatic breast cancer, following promising results from the Phase III DESTINY-Breast 04 and DESTINY-Breast 06 trials, respectively [11,12]. Therefore, accurate and reproducible HER2 assessment, especially at the lower end of the expression spectrum, is essential for effective breast cancer treatment, particularly when considering the potential side effects and costs associated with targeted therapies.

Given the growing recognition of HER2-low and HER2-ultralow breast cancers, several recent studies have evaluated interobserver reproducibility in the assessment of HER2-low, using various versions of the ASCO/CAP guidelines, with inconsistent results [7,[13], [14], [15], [16], [17], [18], [19], [20], [21]]. Most studies had relatively small sample sizes, and only a few specifically included the HER2-ultralow category [13,18,20]. We conducted this study to assess interobserver reproducibility using a large cohort of HER2-negative breast cancer cases in the era of HER2-low and HER2-ultralow breast cancers. The study cohort includes a mix of subspecialized breast pathologists, general surgical pathologists, and a pathologist in fellowship training, with varying years of professional experience, thereby simulating real-world clinical practice.

2. Materials and methods

2.1. Sample collection

After obtaining Institutional Review Board approval, we searched the surgical pathology database at the authors' institution to identify breast cancer cases from 2015 to 2020, in which HER2 testing was performed in-house. A total of 389 cases with HER2 IHC scores of 0, 1+, and 2+, and with archived slides available, were included in the study, comprising 360 biopsies and 29 surgical excisions. Concurrent HER2 ISH tests were performed on all cases according to the institutional protocol [22], and all tests were non-amplified in the cases included.

HER2 IHC testing was performed using an FDA-approved rabbit monoclonal antibody (clone 4B5, prediluted; Ventana, Tucson, AZ) and the Ventana ultraView DAB Detection Kit, following the manufacturer's instructions, on an automated Ventana Benchmark XT platform. HER2 ISH testing was carried out using the FDA-approved INFORM HER2 Dual ISH DNA Probe Cocktail assay (Roche Diagnostics, Indianapolis, IN) on either a VENTANA BenchMark XT or VENTANA BenchMark ULTRA automated slide stainer, as previously described [22].

2.2. Assessment of HER2 immunohistochemistry

To simulate real-world practice, HER2 evaluation was conducted by six board-certified pathologists with varying years of professional experience. The raters included two breast pathologists (with more than 50% of their service assignments dedicated to breast pathology), three surgical pathologists specializing in non-breast fields (with less than 50% of their service assignments in breast pathology), and one surgical pathology fellow/clinical instructor at the time of the study. The years of experience ranged from less than 1 year (n = 1), 1 to 5 years (n = 1), 5 to 10 years (n = 1), 10 to 20 years (n = 2), to more than 20 years (n = 1). Only the research coordinator had access to the pathology reports, and all raters were blinded to the initial HER2 scores.

HER2 evaluation by all raters was performed simultaneously using a multiheaded microscope, the same model as the raters' individual microscopes used in their daily practice. No training session was conducted. The raters were instructed to follow the 2023 ASCO/CAP HER2 testing guidelines [6] and the current CAP Cancer Protocols [23]. In brief, HER2 testing was scored as follows: HER2-null (absent membrane staining), HER2-ultralow (incomplete, faint, or barely perceptible membrane staining in ≤10% of tumor cells), HER2 1+ (incomplete, faint, or barely perceptible membrane staining in >10% of tumor cells), HER2 2+ (weak to moderate complete membrane staining in >10% of tumor cells; moderate to intense but incomplete or basolateral membrane staining; complete and intense circumferential staining in ≤10% of cancer cells; or abundant cytoplasmic staining obscuring evaluation of membrane stain intensity), and HER2 3+ (complete, intense, circumferential membrane staining in >10% of tumor cells). Notably, no case was scored as 3+ by any rater.

No discussion was allowed until the end of each 1-h session, which included approximately 25-30 cases. Best practice recommendations were followed, including initial scanning at 4X, with confirmation at 20-40X for equivocal cases and weak to moderate staining patterns (scores 1+/2+). For cases barely perceptible at 4X and 10X but visible at 20X, 40X was used to help differentiate between HER2-null, HER2-ultralow, and 1+ [6,24]. On average, the time spent on each case was 1.5 to 2 min, including review of the hematoxylin and eosin slide.

2.3. Statistical analysis

The agreement among raters was quantified using Fleiss' kappa, a statistical measure of inter-rater reliability designed for categorical ratings by more than two raters, where 0.01–0.20 is slight, 0.21–0.40 is fair, 0.41–0.60 is moderate, 0.61–0.80 is substantial, and 0.81–1.00 is almost perfect agreement [25]. Fleiss’ kappa adjusts for the agreement expected by chance and is applicable to both nominal and ordinal data. Kappa values were calculated using the kappam. fleiss () function from the irr package in R and interpreted according to the established scale: ≤ 0 indicating no agreement, 0.01–0.20 as slight agreement, 0.21–0.40 as fair, 0.41–0.60 as moderate, 0.61–0.80 as substantial, and 0.81–1.00 as almost perfect agreement [26].

In addition to overall Fleiss' kappa estimates, pairwise kappa values were calculated to assess agreement between individual pathologist pairs. To further evaluate rating consistency, Intraclass Correlation Coefficients (ICCs) were computed using random-effects ANOVA models after numerically recoding the categorical HER2 scores [27]. ICC is a measure of reliability for quantitative or ordinal data and quantifies the proportion of total variance attributable to differences among raters, relative to the total variance observed. ICC values were interpreted as follows: <0.5 indicating poor reliability, 0.5–0.75 as moderate, 0.75–0.9 as good, and >0.9 as excellent reliability. Together, Fleiss’ kappa and ICC analyses provided a comprehensive evaluation of both categorical and numerical agreement across different scoring schemes. Statistical analysis was performed using R software (www.r-project.org).

3. Results

Among the 389 HER2-negative cases evaluated, 157 (40%), 183 (47%), and 49 (13%) were initially scored as 0, 1+, and 2+/ISH-negative, respectively. Notably, among cases initially scored as HER2 0, 54% (84/157) were reclassified as HER2-ultralow by at least one rater, and 22% (34/157) by at least five raters. Furthermore, 34% (53/157) and 15% (23/157) of these IHC score 0 cases were categorized as HER2-low by at least one and at least five raters, respectively. HER2 immunostaining remained interpretable for up to ten years (representative images shown in Fig. 1). Although some degree of deterioration in staining quality may occur over time, the evaluation of this effect was beyond the scope of the present study.

Fig. 1.

Fig. 1

Representative images of HER2 scores with full agreement among all raters. (A) HER2-null; (B) HER2-ultralow; (C) HER2 1+; (D) HER2 2+. All images are shown at original magnification ×400.

3.1. Fleiss’ kappa value estimates

Fleiss' kappa ranges from −1 to +1, with higher values indicating stronger agreement beyond what would be expected by chance. Three different HER2 scoring schemes were used for analysis: a 4-score scheme using the original scores (HER2-null, -ultralow, 1+, and 2+), a 3-score scheme combining HER2 scores 1+ and 2+ as HER2-low (HER2-null, -ultralow, and -low), and a 2-score scheme combining HER2-ultralow and HER2-low as HER2-null and HER2-ultralow/HER2-low (Table 1). The overall Fleiss’ kappa values for all raters across the entire cohort of 389 cases were 0.67, 0.68, and 0.77, respectively, for the three scoring schemes, indicating substantial agreement among the six pathologists.

Table 1.

HER2 scoring schemes for clustering and analysis.

Scheme Clustering Classification
1 Original 4 scores HER2-null, HER2-ultralow, 1+, and 2+
2 3-tier clustering HER2-null, HER2-ultralow, and HER2-low (1+/2+ combined)
3 2-tier clustering HER2-null, HER2-ultralow/HER2-low (combined)

Fleiss' kappa values were also estimated for each pair of raters to evaluate their agreement. Of the 15 pairs, 10 showed substantial agreement using the original 4-score scheme (κ = 0.61–0.80), while the remaining 5 pairs demonstrated moderate agreement (κ = 0.41–0.60) (Table 2). Given that HER2-low breast cancers are clinically managed similarly, we next estimated Fleiss' kappa values using the 3-score scheme. A similar, but slightly improved trend was observed: one pair reached almost perfect agreement (κ = 0.81–1.00), 9 pairs showed substantial agreement, and 5 pairs showed moderate agreement (Table 2). Moreover, considering that trastuzumab deruxtecan has shown similar survival benefits in patients with HER2-low and HER2-ultralow metastatic breast cancer [12], further analysis using the 2-score scheme—combining HER2-ultralow and HER2-low—resulted in a greater improvement in Fleiss’ kappa values: 4 pairs with almost perfect agreement and 11 pairs with substantial agreement (Table 2). Thus, clustering the scores into fewer categories based on similar treatment modalities improves inter-rater agreement.

Table 2.

Fleiss’ kappa values for each pair of raters.

Rater 4-tier: null vs. ultralow vs. 1+ vs. 2+
3-tier: null vs. ultralow vs. low
2-tier: null vs. ultralow/low
2 3 4 5 6 2 3 4 5 6 2 3 4 5 6
1 0.61 0.74 0.71 0.68 0.54 0.64 0.77 0.73 0.71 0.54 0.71 0.78 0.81 0.72 0.66
2 0.79 0.71 0.76 0.57 0.80 0.73 0.77 0.57 0.88 0.87 0.77 0.78
3 0.79 0.75 0.58 0.84 0.75 0.57 0.90 0.79 0.70
4 0.69 0.58 0.71 0.58 0.78 0.69
5 0.59 0.56 0.71

It is of particular interest that all instances of moderate agreement consistently involved a single rater (Rater #6). To investigate this further, we asked whether this rater consistently provided lower or random ratings compared to the other evaluators. Indeed, Rater 6 generally rated lower than the remaining five raters (data not shown). This observation was further confirmed by two one-sample t-tests. First, for each patient, the score was calculated as the rating of Rater 6 minus the average rating of the other five raters. The average score across all samples was −0.115, and the p-value for the average score being less than 0 was 1.12 × 10−6, indicating that the average score was significantly lower than 0. Second, for each patient, the score was calculated as the difference between the number of raters who rated lower than Rater 6 and the number of raters who rated higher than Rater 6. The average score across all samples was −0.545, and the p-value for the average score being less than 0 was 1.26 × 10−6, further indicating that the average score was significantly lower than 0. Therefore, Rater 6 consistently gave lower scores compared to the other raters, suggesting a higher threshold for scoring.

3.2. Intraclass correlation

To further assess inter-rater consistency, we calculated the intraclass correlation coefficient (ICC). The ICC ranges from 0 (no reliability) to 1 (perfect reliability among raters) [27]. As the ICC is derived from random-effects ANOVA models, HER2 scores were numerically recoded prior to analysis: 1–4 for the 4-score scheme, 1–3 for the 3-score scheme, and 1–2 for the 2-score scheme.

ICC values were calculated using the icc () (irr package) and ICC() (psych package) functions in R. The one-way model assumes random sampling of samples only, whereas the two-way model assumes random sampling of both samples and raters. To that end, the irr::icc () and psych::ICC() functions yielded similar results (Table 3). All ICCs ranged from 0.75 to 0.90, demonstrating good reliability. These findings align with the observations from the Fleiss’ kappa values. A minor decline in ICC values was noted with fewer HER2 categories, suggesting that simplifying the classification scheme may modestly reduce reliability.

Table 3.

Estimated ICC values for consistency.

Model icc() - irr package
null vs. ultralow vs. 1+ vs. 2+
null vs. ultralow vs. low
null vs. ultralow/low
ICC F-test P 95% CI ICC F-test P 95% CI ICC F-test P 95% CI
One-way 0.84 33.3 <10−16 0.82, 0.86 0.83 29.7 <10−16 0.80, 0.85 0.77 21 <10−16 0.74, 0.80
Two-way 0.85 34.2 <10−16 0.83, 0.87 0.83 31 <10−16 0.81, 0.86 0.77 21.2 <10−16 0.74, 0.80

ICC() - psych package

One-way 0.85 34 <10−16 0.83, 0.87 0.83 31 <10−16 0.81, 0.85 0.77 21 <10−16 0.74, 0.80
Two-way 0.84 34 <10−16 0.82, 0.86 0.83 31 <10−16 0.80, 0.85 0.77 21 <10−16 0.74, 0.80

CI, confidence interval.

3.3. Descriptive analysis

The distribution of HER2 scores was subsequently analyzed. Using the original 4-score scheme, 206 (53.0%), 172 (44.2%), 10 (2.6%), and 1 (0.3%) cases received 1, 2, 3, and 4 scores, respectively (Table 4). Among all cases, 206 (53.0%) exhibited complete agreement across all six raters. In 94 (24.2%) and 49 (12.6%) cases, two scores were given, with one assigned by five or four raters, respectively. Overall, 89.7% (349/389) of cases showed full or near-complete agreement among raters (i.e., two scores, one assigned by ≥ 4 raters).

Table 4.

Distribution of scores.

No. Scores No. Cases Cases scored by 6 raters Cases scored by 5 raters Cases scored by 4 raters Cases scored by 3 raters Cases scored by 2 raters
4-tier: null vs. ultralow vs. 1+ vs. 2+
1 206 (53.0%) 206 (53.0%)
2 172 (44.2%) 94 (24.2%) 49 (12.6%) 29 (7.5%)
3 10 (2.6%) 3 (0.8%) 6 (1.5%) 1 (0.3%)
4 1 (0.3%) 1 (0.3%)
3-tier: null vs. ultralow vs. low
1 238 (61.2%) 238 (61.2%)
2 142 (36.5%) 83 (21.3%) 36 (9.3%) 23 (5.9%)
3 9 (2.3%) 4 (1.0%) 4 (1.0%) 1 (0.3%)
2-tier: null vs. ultralow/low
1 338 (86.9%) 338 (86.9%)
2 51 (13.1%) 30 (7.7%) 10 (2.6%) 11 (2.8%)

Under the 3-score scheme, 238 (61.2%), 142 (36.5%), and 9 (2.3%) cases exhibited 1, 2, and 3 distinct scores, respectively. Among these, 91.8% (357/389) of cases demonstrated either complete agreement across all six raters or two scores, one assigned by at least four raters. As expected, agreement further improved with the 2-score scheme: 338 (86.9%) and 51 (13.1%) cases received 1 and 2 distinct scores, respectively, and 97.2% (378/389) of cases exhibited either full agreement or near-consensus (≥4 raters giving the same score) (Table 4). In line with these findings, the greatest inter-rater consensus was observed in the combined HER2-ultralow/-low group. Of the 356 cases identified as HER2-ultralow, 1+, or 2+ by at least one rater, 93.3%, 90.7%, and 85.7% reached agreement among at least four, five, and all six raters, respectively (Table 5).

Table 5.

HER2 score classification across increasing rater agreement thresholds.

HER2 score Any rater (≥1)a At least 4 ratersb At least 5 ratersb All 6 ratersb
HER2-null 84 46 (54.8%) 45 (53.6%) 33 (39.3%)
HER2-ultralow 184 76 (41.3%) 55 (29.9%) 33 (17.9%)
HER2 1+ 255 195 (76.5%) 170 (66.7%) 115 (45.1%)
HER2 2+ 61 35 (57.4%) 30 (49.2%) 25 (41.0%)
HER2-low 281 239 (85.1%) 221 (78.6%) 172 (61.2%)
HER2-ultralow/-low 356 332 (93.3%) 323 (90.7%) 305 (85.7%)
a

Any rater indicates the number of cases assigned to a given category by at least one individual rater. This reflects any rater–case assignment and does not imply consensus.

b

At least n raters indicate the number of cases assigned to a given category by at least the specified number of raters, reflecting progressively greater interobserver agreement.

The distribution of HER2 scores across all 389 cases, using the different scoring schemes, is illustrated in Fig. 2.

Fig. 2.

Fig. 2

Distribution of HER2 scores across the 389-case cohort using different scoring schemes. Top panel: 4-tier scheme (HER2-null, HER2-ultralow, 1+, 2+). Middle panel: 3-tier scheme (HER2-null, HER2-ultralow, HER2-low [1+/2+ combined]). Bottom panel: 2-tier scheme (HER2-null, HER2-ultralow/-low [ultralow/low combined]). Each sample on the X-axis represents a case, with stacked bars showing the proportion of raters assigning each score. This visualization highlights interobserver variability and demonstrates how simplifying the scoring scheme increases consensus. UL, ultralow.

4. Discussion

HER2 protein expression in breast cancer has historically been categorized as 0, 1+, 2+, or 3+ over the past two decades. This scoring system was originally established to identify HER2-positive tumors (score 3+ or 2+/amplified) eligible for HER2-targeted therapies [2]. However, HER2 expression exists along a continuum rather than a binary positive–negative spectrum. HER2-low status is observed in approximately 40–60% of all breast cancers [8]. Moreover, the reported incidence of HER2-ultralow breast cancer varies across studies depending on exclusion criteria and patient demographics, ranging from 10% to 30% [19,[28], [29], [30], [31]]. In the present cohort, 54% and 22% of cases initially scored as HER2 IHC score 0 were reclassified as HER2-ultralow by at least one and at least five raters, respectively, consistent with recent findings [31]. As HER2-targeted therapies expand to include HER2-ultralow tumors, accurate classification is increasingly critical for guiding precision treatment decisions.

The concordance of HER2 IHC scoring in the low range (0 and 1+) among pathologists has been reported to be suboptimal [7]. This is largely attributable to the original design of the scoring system, which aimed primarily to reliably identify HER2 score 3+ cases. For the past two decades, there was little incentive for pathologists to distinguish between HER2 scores of 0 and 1+, as this distinction had no impact on clinical decision-making. Only recently, with the emerging clinical relevance of HER2-low and HER2-ultralow categories, has accurate discrimination between these lower scores gained significance.

Few studies have evaluated the interobserver reproducibility of HER2 scoring incorporating the HER2-ultralow category, and their findings have been variable. In one study evaluating 50 previously classified HER2 IHC score 0 cases by 36 pathologists from multiple centers, Fleiss’ kappa values were 0.344 and 0.230 when using three (HER2 score 0, 1+, 2+) and four (HER2-null, -ultralow, 1+, 2+) scoring categories, respectively [18]. In contrast, another recent study of 247 cases reported substantial agreement (κ = 0.66–0.70) among three experienced breast pathologists when assessing the lower end of the HER2 expression spectrum (null, ultralow, 1+; n = 147) [20]. These findings are consistent with our observations, where an overall substantial agreement was achieved, even when applying the four-tiered scoring system. Importantly, scoring between glass slides and digital images was largely concordant, although digital image evaluation may be slightly more sensitive at very low levels of immunostaining [20].

This study offers several noteworthy highlights. First, the findings reflect real-world practice within a large academic medical center, where the raters included a diverse group of pathologists - those specializing in breast pathology, others primarily focused on different subspecialties and less frequently involved in breast cases, as well as surgical pathology fellows in training. The participants also represented a broad range of experience levels, from less than one year to over twenty years in practice, and no dedicated training session was conducted prior to the evaluation. Consequently, the results are considered to be more representative and generalizable than studies limited to “expert-only” participants.

Second, experience level and training background did not appear to significantly influence scoring concordance among pathologists who routinely evaluate HER2. All staff pathologists had received training and practiced at different institutions prior to joining the current center, yet their performance was comparable. Notably, the least experienced participant - a pathology fellow in training (Rater #5) - performed at a level similar to that of most other raters.

Third, individual raters’ pre-existing interpretive thresholds likely contributed to the major discrepancies observed (e.g., Rater #6 versus others). This may partly explain the limited improvement in reproducibility reported after brief training sessions [13], as such sessions are unlikely to counterbalance years of ingrained interpretive experience. Conversely, the routine practice at our institution of reaching an internal consensus for final HER2 score assignment in challenging cases likely played a key role in the overall superior interobserver agreement observed in this study compared with others. This approach has been recognized as a best practice [24] and likely contributes to the more closely aligned interpretive thresholds developed among our group over time.

Moreover, studies such as ours provide a practical framework for identifying systematic interpretive differences that arise from individual experience and learned behaviors. Rather than relying on feedback from isolated cases, comparative data across multiple raters enable pathologists to recognize consistent differences between their own scoring patterns and those of their peers, facilitating deliberate calibration of their interpretive thresholds. The insights gained from this study have led to more frequent internal consensus discussions for borderline cases at our institution's daily breast pathology consensus conference. Thus, the value of this study extends beyond documenting real-world interobserver variability by demonstrating a practical approach to improving consistency in HER2 IHC interpretation.

Lastly, lower interobserver consistency was most often observed in cases at the interface between HER2-null and HER2-ultralow, where faint, incomplete staining can be subtle and distinguishing true membranous staining from non-specific cytoplasmic reactivity is challenging. Similar difficulties were noted in cases bordering the HER2-ultralow and 1+ categories. Accordingly, simplifying the classification by consolidating scores into fewer groups can improve reproducibility. Although quantitative assessment of HER2 expression may serve as a useful adjunct for subclassifying HER2-unamplified breast cancers, it cannot fully address intratumoral heterogeneity—reported in more than 40% of HER2-low breast cancers [32] and considered nearly universal in HER2-ultralow tumors.

In summary, substantial overall agreement was achieved in assessing a large cohort of breast cancer cases at the lower end of the HER2 expression spectrum. This study reflects real-world diagnostic practice and demonstrates that simplifying the HER2 scoring scheme enhances interobserver concordance among pathologists. Despite the nuanced distinction between HER2-low and HER2-ultralow categories, IHC remains a reliable and reproducible method for evaluating HER2 expression in breast cancer, even in the era of expanded HER2-targeted therapies. Achieving consensus on final HER2 score designation in borderline cases should be encouraged to ensure optimal treatment selection.

CRediT authorship contribution statement

Xiao Huang: Writing – review & editing, Data curation, Conceptualization. Sarah Anderson: Writing – review & editing, Data curation. Valeria Dal Zotto: Writing – review & editing, Data curation. Shuko Harada: Writing – review & editing, Data curation. Andrea G Kahn: Writing – review & editing, Data curation. Meiling Zhou: Writing – review & editing, Formal analysis. Kui Zhang: Writing – review & editing, Formal analysis. Kanako Okamoto: Writing – review & editing, Data curation. Shi Wei: Writing – review & editing, Writing – original draft, Supervision, Project administration, Methodology, Formal analysis, Data curation, Conceptualization.

Data availability

Data available within the article.

Ethical approval and consent to participate

This research was approved by the Institutional Review Board at the University of Alabama at Birmingham (IRB-090814006). Informed consent was not required per IRB given the retrospective, non-interventional nature of this research. Research was performed in accordance with the Declaration of Helsinki.

Consent for publication

Not required per IRB.

Funding

None.

Declaration of Competing interests

SW is a consultant on the Breast Pathology Faculty Advisory Board for Daiichi Sankyo Inc and AstraZeneca (not relevant to this study). All other authors declare they have no financial interests.

Acknowledgements

None.

References

  • 1.Marra A., Chandarlapaty S., Modi S. Management of patients with advanced-stage HER2-positive breast cancer: current evidence and future perspectives. Nat Rev Clin Oncol. 2024;21:185–202. doi: 10.1038/s41571-023-00849-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2.Slamon D.J., Leyland-Jones B., Shak S., Fuchs H., Paton V., Bajamonde A., et al. Use of chemotherapy plus a monoclonal antibody against HER2 for metastatic breast cancer that overexpresses HER2. N Engl J Med. 2001;344:783–792. doi: 10.1056/NEJM200103153441101. [DOI] [PubMed] [Google Scholar]
  • 3.Wolff A.C., Hammond M.E., Schwartz J.N., Hagerty K.L., Allred D.C., Cote R.J., et al. American society of clinical oncology/college of American pathologists guideline recommendations for human epidermal growth factor receptor 2 testing in breast cancer. J Clin Oncol. 2007;25:118–145. doi: 10.1200/JCO.2006.09.2775. [DOI] [PubMed] [Google Scholar]
  • 4.Wolff A.C., Hammond M.E., Hicks D.G., Dowsett M., McShane L.M., Allison K.H., et al. Recommendations for human epidermal growth factor receptor 2 testing in breast cancer: American society of clinical oncology/college of American pathologists clinical practice guideline update. J Clin Oncol. 2013;31:3997–4013. doi: 10.1200/JCO.2013.50.9984. [DOI] [PubMed] [Google Scholar]
  • 5.Wolff A.C., Hammond M.E.H., Allison K.H., Harvey B.E., Mangu P.B., Bartlett J.M.S., et al. Human epidermal growth factor receptor 2 testing in breast cancer: American society of clinical oncology/college of American pathologists clinical practice guideline focused update. J Clin Oncol. 2018;36:2105–2122. doi: 10.1200/JCO.2018.77.8738. [DOI] [PubMed] [Google Scholar]
  • 6.Wolff A.C., Somerfield M.R., Dowsett M., Hammond M.E.H., Hayes D.F., McShane L.M., et al. Human epidermal growth factor receptor 2 testing in breast cancer. Arch Pathol Lab Med. 2023;147:993–1000. doi: 10.5858/arpa.2023-0950-SA. [DOI] [PubMed] [Google Scholar]
  • 7.Fernandez A.I., Liu M., Bellizzi A., Brock J., Fadare O., Hanley K., et al. Examination of low ERBB2 protein expression in breast cancer tissue. JAMA Oncol. 2022;8:1–4. doi: 10.1001/jamaoncol.2021.7239. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.Tarantino P., Hamilton E., Tolaney S.M., Cortes J., Morganti S., Ferraro E., et al. HER2-Low breast cancer: pathological and clinical landscape. J Clin Oncol. 2020;38:1951–1962. doi: 10.1200/JCO.19.02488. [DOI] [PubMed] [Google Scholar]
  • 9.Tarantino P., Viale G., Press M.F., Hu X., Penault-Llorca F., Bardia A., et al. vol. 34. Ann Oncol; 2023. pp. 645–659. (ESMO expert consensus statements (ECS) on the definition, diagnosis, and management of HER2-low breast cancer). [DOI] [PubMed] [Google Scholar]
  • 10.Nicolò E., Boscolo Bielo L., Curigliano G., Tarantino P. The HER2-low revolution in breast oncology: steps forward and emerging challenges. Ther Adv Med Oncol. 2023;15 doi: 10.1177/17588359231152842. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11.Modi S., Jacot W., Yamashita T., Sohn J., Vidal M., Tokunaga E., et al. Trastuzumab deruxtecan in previously treated HER2-Low advanced breast cancer. N Engl J Med. 2022;387:9–20. doi: 10.1056/NEJMoa2203690. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Bardia A., Hu X., Dent R., Yonemori K., Barrios C.H., O'Shaughnessy J.A., et al. Trastuzumab deruxtecan after endocrine therapy in metastatic breast cancer. N Engl J Med. 2024;391:2110–2122. doi: 10.1056/NEJMoa2407086. [DOI] [PubMed] [Google Scholar]
  • 13.Baez-Navarro X., van Bockstal M.R., Nawawi D., Broeckx G., Colpaert C., Doebar S.C., et al. Interobserver variation in the assessment of immunohistochemistry expression levels in HER2-Negative breast cancer: can we improve the identification of low levels of HER2 expression by adjusting the criteria? An international interobserver study. Mod Pathol. 2023;36 doi: 10.1016/j.modpat.2022.100009. [DOI] [PubMed] [Google Scholar]
  • 14.Karakas C., Tyburski H., Turner B.M., Wang X., Schiffhauer L.M., Katerji H., et al. Interobserver and interantibody reproducibility of HER2 immunohistochemical scoring in an enriched HER2-Low-Expressing breast cancer cohort. Am J Clin Pathol. 2023;159:484–491. doi: 10.1093/ajcp/aqac184. [DOI] [PubMed] [Google Scholar]
  • 15.Schettini F., Chic N., Brasó-Maristany F., Paré L., Pascual T., Conte B., et al. Clinical, pathological, and PAM50 gene expression features of HER2-low breast cancer. NPJ breast cancer. 2021;7:1. doi: 10.1038/s41523-020-00208-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Sun H., Kang E.Y., Chen H., Sweeney K.J., Suchko M., Wu Y., et al. Immunohistochemical assessment of HER2 low breast cancer: interobserver reproducibility and correlation with digital image analysis. Breast Cancer Res Treat. 2024;205:403–411. doi: 10.1007/s10549-024-07256-3. [DOI] [PubMed] [Google Scholar]
  • 17.Zaakouk M., Quinn C., Provenzano E., Boyd C., Callagy G., Elsheikh S., et al. Concordance of HER2-low scoring in breast carcinoma among expert pathologists in the United Kingdom and the Republic of Ireland -on behalf of the UK national coordinating committee for breast pathology. Breast. 2023;70:82–91. doi: 10.1016/j.breast.2023.06.005. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Wu S., Shang J., Li Z., Liu H., Xu X., Zhang Z., et al. Interobserver consistency and diagnostic challenges in HER2-ultralow breast cancer: a multicenter study. ESMO Open. 2025;10 doi: 10.1016/j.esmoop.2024.104127. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.Lv H., Yue J., Zhang Q., Xu F., Gao P., Yang H., et al. Prevalence and concordance of HER2-low and HER2-ultralow status between historical and rescored results in a multicentre study of breast cancer patients in China. Breast Cancer Res. 2025;27:45. doi: 10.1186/s13058-025-02001-0. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.Xiao A., Vohra P., Chen Y.Y., Ung L., Kim M.O., Geradts J. Comparative study of intra- and inter-observer variability in manual scoring of HER2 immunohistochemical stains on glass slides versus paired digital images with emphasis on the low end of the expression spectrum. Hum Pathol. 2025;161 doi: 10.1016/j.humpath.2025.105860. [DOI] [PubMed] [Google Scholar]
  • 21.Robbins C.J., Fernandez A.I., Han G., Wong S., Harigopal M., Podoll M., et al. Multi-institutional assessment of pathologist scoring HER2 immunohistochemistry. Mod Pathol. 2023;36 doi: 10.1016/j.modpat.2022.100032. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Memon R., Prieto Granada CN., Harada S., Winokur T., Reddy V., Kahn A.G., et al. Discordance between immunohistochemistry and in situ hybridization to detect HER2 overexpression/gene amplification in breast cancer in the Modern Age: a single institution experience and pooled literature review study. Clin Breast Cancer. 2022;22:e123–e133. doi: 10.1016/j.clbc.2021.05.004. [DOI] [PubMed] [Google Scholar]
  • 23.Allison K.H., Krishnamurti U. College of American Pathologists; 2025. Template for reporting results of biomarker testing of specimens from patients with carcinoma of the breast.https://documents.cap.org/documents/New-Cancer-Protocols-June-2025/Breast.Bmk_1.6.1.0.-REL_CAPCP.pdf?_gl=1*jy8asy*_ga*MTU5MTYxODgzMi4xNzE5Nzc0NDQ0*_ga_97ZFJSQQ0X*czE3NTI1MjcxMjIkbzM4JGcxJHQxNzUyNTI3MTM5JGo0MyRsMCRoMA Version: 1.6.1.0 ed. [Google Scholar]
  • 24.Tozbikian G., Bui M.M., Hicks D.G., Jaffer S., Khoury T., Wen H.Y., et al. Best practices for achieving consensus in HER2-low expression in breast cancer: current perspectives from practising pathologists. Histopathology. 2024;85:489–502. doi: 10.1111/his.15275. [DOI] [PubMed] [Google Scholar]
  • 25.Fleiss J.L. Measuring nominal scale agreement among many raters. Psychol Bull. 1971;76(5):378–382. [Google Scholar]
  • 26.McHugh M.L. Interrater reliability: the kappa statistic. Biochem Med. 2012;22:276–282. [PMC free article] [PubMed] [Google Scholar]
  • 27.Shrout P.E., Fleiss J.L. Intraclass correlations: uses in assessing rater reliability. Psychol Bull. 1979;86:420–428. doi: 10.1037//0033-2909.86.2.420. [DOI] [PubMed] [Google Scholar]
  • 28.Hu Y., Jones D., Zhao W., Tozbikian G., Parwani A.V., Li Z. Incidence, clinicopathologic features, and Follow-Up results of human epidermal growth factor receptor-2-Ultralow breast carcinoma. Mod Pathol. 2025;38 doi: 10.1016/j.modpat.2025.100783. [DOI] [PubMed] [Google Scholar]
  • 29.Guan F., Ju X., Chen L., Ren J., Ke X., Luo B., et al. Comparison of clinicopathological characteristics, efficacy of neoadjuvant therapy, and prognosis in HER2-low and HER2-ultralow breast cancer. Diagn Pathol. 2024;19:131. doi: 10.1186/s13000-024-01557-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30.Chen Z., Jia H., Zhang H., Chen L., Zhao P., Zhao J., et al. Is HER2 ultra-low breast cancer different from HER2 null or HER2 low breast cancer? A study of 1363 patients. Breast Cancer Res Treat. 2023;202:313–323. doi: 10.1007/s10549-023-07079-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31.Morrar D., Brogi E., Schwartz C.J., Pareja F., Wen H.Y., Ross D.S. Human epidermal growth factor receptor 2 (HER2)-ultralow breast cancer: incidence, clinicopathologic features, and need for refined scoring system. Mod Pathol. 2025;38 doi: 10.1016/j.modpat.2025.100901. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32.Zhang H., Katerji H., Turner B.M., Audeh W., Hicks D.G. HER2-low breast cancers: incidence, HER2 staining patterns, clinicopathologic features, MammaPrint and BluePrint genomic profiles. Mod Pathol. 2022;35:1075–1082. doi: 10.1038/s41379-022-01019-5. [DOI] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Data Availability Statement

Data available within the article.


Articles from The Breast : Official Journal of the European Society of Mastology are provided here courtesy of Elsevier

RESOURCES