Nomogram based on deep learning of mammography for the prediction of HER2 expression in breast cancer
Highlight box
Key findings
• Two nomograms integrating deep learning (DL)-derived imaging features and clinical variables were developed to predict human epidermal growth factor receptor 2 (HER2) status. The models effectively distinguished not only HER2-positive from HER2-negative tumors but also HER2-low from HER2-zero subtypes. At the same time, future efforts should focus on prospective validation and integration of the models into clinical workflows.
What is known and what is new?
• HER2 status is critical for guiding treatment decisions in breast cancer, and DL-based imaging models have been explored for binary HER2 classification. However, relatively limited studies have explored further stratification of HER2-negative breast cancer.
• This study established DL-based nomograms that enable further differentiation between HER2-low and HER2-zero tumors, a distinction increasingly relevant for antibody-drug conjugate (ADC) therapy. In addition, full-field mammographic images were utilized, eliminating the need for manual tumor segmentation.
What is the implication, and what should change now?
• The proposed models provide a non-invasive approach for preoperative HER2 assessment, which may support treatment planning and patient stratification. Meanwhile, identifying HER2-low tumors may help select patients who could benefit from ADC therapy. Following further prospective validation, this model could be integrated into routine clinical workflows.
Introduction
The incidence of breast cancer, the most prevalent malignant tumor in women, is steadily rising. At first, the approach to treating breast cancer was determined by the presence of hormone receptor (HR) and the Ki67 proliferation index (1). Subsequently, there has been a growing interest in the predictive significance of human epidermal growth factor receptor 2 (HER2). The HER2 status in breast cancer has both independent prognostic and therapeutic predictive significance (2,3). In 2018, the American Society of Clinical Oncology (ASCO) and the College of American Pathologists (CAP) classified HER2 into two categories: positive and negative. Molecular assessment of HER2 is therefore based on a binary determination of positivity or negativity (4). However, in the absence of fluorescence in situ hybridization (FISH) gene amplification, HER2-negative breast cancer can be further classified into HER2-low and HER2-zero expression subtypes. HER2 positivity is defined by an immunohistochemistry (IHC) score of 3+ or gene amplification detected by FISH, whereas an IHC score of 0 or 1+ indicates HER2 negativity. In the absence of FISH gene amplification, HER2 status is considered negative. Compared with HER2-negative breast cancer, HER2 overexpression is generally associated with more aggressive tumor behavior, higher risk of recurrence, and reduced overall survival (OS) (5-7). Research has demonstrated that conventional HER2-targeted treatments, such as trastuzumab, pertuzumab, and lapatinib, can greatly enhance the life expectancy of patients with HER2-positive breast cancer (8-13). Multiple studies have demonstrated that HER2-targeted therapies, such as the antibody-drug conjugate (ADC) trastuzumab deruxtecan (T-DXd), significantly improve progression-free survival (PFS) and OS in patients with HER2-low metastatic breast cancer. In contrast, patients with HER2-zero tumors derive little or no clinically meaningful benefit from these ADCs (14-18). Therefore, accurate assessment of HER2 status at the initial stage of treatment is crucial for identifying patients suitable for HER2-targeted interventions.
Currently, most studies focus on predicting HER2-positive and HER2-negative status in breast cancer, whereas few have investigated the prediction of HER2-low and HER2-zero expression (19-23). In the current era of widespread use of ADCs for patients with HER2-low expression, the traditional binary classification of HER2 status has become insufficient for clinical decision-making. At the same time, HER2 expression is assessed through the use of IHC and/or in situ hybridization (ISH) techniques on tissue samples, which can be considered somewhat invasive. Advances in science and technology have increasingly enabled the application of artificial intelligence to efficiently analyze large-scale medical data. This study integrates widely used mammographic X-ray imaging and clinical parameters with artificial intelligence algorithms to develop two independent binary nomograms for identifying three HER2 breast cancer subtypes (positive, low, and zero expression). The proposed approach enables noninvasive breast cancer detection and early assessment of HER2 status prior to pathological confirmation, thereby serving as a complementary tool to support clinical decision-making. By employing a comprehensive and multifaceted approach, we seek to establish a robust scientific basis for evaluating the potential benefits of targeted therapies and optimizing treatment strategies. Furthermore, the imaging model does not require manual tumor segmentation, thereby substantially reducing human intervention and yielding significant savings in both labor and economic costs. We present this article in accordance with the TRIPOD reporting checklist (available at https://gs.amegroups.com/article/view/10.21037/gs-2025-aw-467/rc).
Methods
Patients
The study was conducted in accordance with the Declaration of Helsinki and its subsequent amendments. The study was approved by the Ethics Committee of Affiliated Hospital of Jining Medical University (ethics approval No. 2021C019) and individual consent for this retrospective analysis was waived.
The inclusion criteria of this multi-center observational study are listed as follows.
Inclusion criteria were as follows: (I) postoperative pathological confirmation of non-specific invasive breast cancer; (II) preoperative mammography examination performed; (III) postoperative IHC and/or ISH conducted; (IV) complete clinical data available.
Exclusion criteria included: (I) receipt of neoadjuvant chemotherapy prior to surgery; (II) poor-quality mammography images; (III) history of prior tumors or other malignancies.
Following the above inclusion criteria, 406 breast cancer patients who underwent mammography at the main campus of Jining Medical University Affiliated Hospital between January 2019 and June 2024 were assigned to the training and internal validation sets. An independent external validation set comprised 137 breast cancer patients who underwent mammography at Jining Medical University Affiliated Hospital (Taibai Lake Campus) from January 2023 to June 2024. Patient data were retrieved from the Picture Archiving and Communication System (PACS). Data represent biological replicates, derived from independent patient imaging samples. The training and internal validation sets were randomly assigned at a 7:3 ratio (Figure 1).
HER2 assessment
Pathologically, the expression levels of HER2 in postoperative specimens are classified as IHC 0, 1+, 2+, or 3+, strictly following the ASCO/CAP guidelines (24). Subsequently, FISH analysis was performed on samples with IHC status of 2+. According to the combined results of IHC and FISH, patients with IHC 1+ or IHC 2+ status and FISH-negative were classified as HER2-low, while patients with IHC 0 status were classified as HER2-zero. FISH-positive IHC 3+ or IHC 2+ patients were classified as HER2-positive. Cases with indeterminate IHC findings or IHC 2+ status without available FISH results were excluded from the analysis.
Imaging examinations
In the main campus of Jining Medical University Affiliated Hospital, mammography was performed using a Siemens Mammomat digital mammography system (Siemens Healthineers, Germany). Standard mediolateral oblique (MLO) and craniocaudal (CC) views were obtained with a tube voltage of 23–35 kV and a current of 40–80 mAs under automatic exposure control. A breast compression plate was used to maintain a compression thickness between 25 and 40 mm.
In Jining Medical University Affiliated Hospital (Taibai Lake Campus), mammography was performed using a Hologic ASY-00676 digital mammography system (Hologic Inc., USA). Standard MLO and CC views were acquired with a tube voltage of 30–35 kV and a current of 50–80 mAs under automatic exposure control. Breast compression was applied to achieve a thickness between 25 and 40 mm.
Image preprocessing and deep learning (DL) feature extraction
This study utilized whole-breast images without any manual annotation of regions of interest (ROIs). The MLO and CC images were retrieved from the PACS system in Digital Imaging and Communication in Medicine (DICOM) format and subsequently converted to JPG format. Interfering information in the original DICOM images, including machine borders, orientation labels, and blank padding, was removed to facilitate subsequent analysis. Mammography images were selected using a pre-trained ResNet-50 model on the ImageNet dataset to extract features. The MLO and CC views were processed separately through two independent ResNet-50 branches. Features were extracted from the average pooling (AvgPool) layer, followed by dimensionality reduction using principal component analysis (PCA). The reduced feature representations were then concatenated for subsequent analysis.
Feature selection and signature construction
The DL features extracted in the previous step were subjected to feature screening. First, to identify highly redundant features, the Pearson correlation coefficient was calculated to evaluate pairwise feature correlations. For any pair of features with a correlation coefficient greater than 0.9, only one feature was retained to reduce redundancy and enhance feature robustness. Subsequently, the least absolute shrinkage and selection operator (LASSO) regression was applied for feature selection, with 10-fold cross-validation used to improve the reliability of selected DL features. Finally, a prediction model was constructed based on the extracted features using the Light Gradient Boosting Machine (LightGBM) algorithm. During model training, the built-in class weight adjustment strategy of LightGBM was employed to mitigate the effects of class imbalance.
Model establishment and evaluation
Breast cancer subtypes were associated with gland composition, calcification, age, maximum tumor diameter, and cancer antigen 15-3 (CA15-3) and carcinoembryonic antigen (CEA) levels. To identify clinically significant variables, univariate logistic regression (variables with P<0.05 were considered potential predictors and subsequently included in the multivariate analysis) was first performed in the training cohort. Variables with statistical significance (P<0.05) were then included in multivariate logistic regression to identify independent predictors, which were used to construct the clinical model. Subsequently, a combined model integrating both independent clinical predictors and DL features (λ ≠ 0) was established to assess whether this integration improved predictive performance. Three models—the clinical model, the mammography model, and the nomogram—were evaluated for HER2 classification using receiver operating characteristic (ROC) curve analysis. Model performance was assessed in terms of the area under the ROC curve (AUC), sensitivity, and specificity. The clinical model, mammography model, and nomogram AUCs were compared using the DeLong test. Decision curve analysis (DCA) was performed to evaluate the clinical net benefit of each model across a range of threshold probabilities. Calibration curves were plotted to assess the agreement between the predicted probabilities and the observed outcomes.
Statistical analysis
This study included only patients with complete clinical data and a complete-case analysis was performed without any data imputation. All clinical variables incorporated in the analysis—including maximum tumor diameter, CEA, CA15-3, age, gland composition, and calcification—had no missing values. Clinical characteristics were compared between HER2-negative and HER2-positive patients to identify statistically significant differences between cases with HER2-low and HER2-zero subgroups. The mean and standard deviation were expressed as quantitative variables. Frequencies and proportions were summarized as categorical variables. The Mann-Whitney U test was used for non-normal quantitative variables, while the Student’s t-test was used for normal quantitative variables. The Chi-square (χ²) test was used to analyze categorical variables. Univariate and multivariate logistic regression models calculated the odds ratios (ORs) and 95% confidence intervals (CIs). DL feature selection used LASSO regression. The ROC curve and DCA were used to evaluate the model’s discrimination and clinical applicability, respectively. The calibration curve compared model predictions to observed results. All statistical analyses were performed using Statsmodels (version 0.11.1) and R software (version 4.1.3; R Foundation for Statistical Computing, Vienna, Austria). The scikit-learn package (version 1.1.3) was used for machine learning analyses in Python. A two-sided P value <0.05 was considered statistically significant.
Reproducibility statement
All feature extraction, feature selection, and model training procedures were performed exclusively on the training set. The internal and external validation sets were used solely for performance evaluation.
Results
Clinical information of patients
From January 2019 to June 2024, a total of 543 patients were enrolled and subsequently divided into a training set, an internal validation set, and an external validation set. All 543 patients were used to construct Nomogram (A), which was developed to discriminate between HER2-negative and HER2-positive status. For the construction of Nomogram (a), aimed at distinguishing HER2-zero from HER2-low expression, 418 HER2-negative patients were included (Tables 1,2).
Table 1
| Characteristics | Sort | All | HER2-negative | HER2-positive | P |
|---|---|---|---|---|---|
| Gland composition | 25–50% | 290 (53.407) | 211 (50.478) | 79 (63.200) | 0.09 |
| 51–75% | 191 (35.175) | 156 (37.321) | 35 (28.000) | ||
| <25% | 54 (9.945) | 45 (10.766) | 9 (7.200) | ||
| >75% | 8 (1.473) | 6 (1.435) | 2 (1.600) | ||
| Calcification | Contain | 384 (70.718) | 296 (70.813) | 88 (70.400) | 0.93 |
| Not contain | 159 (29.282) | 122 (29.187) | 37 (29.600) | ||
| Age, years | – | 53.000 [47.000, 60.000] | 53.000 [46.000, 61.000] | 54.000 [49.000, 60.000] | 0.30 |
| Maximum diameter, cm | – | 2.300 [1.700, 3.200] | 2.200 [1.700, 2.800] | 3.000 [2.200, 4.200] | <0.001*** |
| CA153, U/mL | – | 8.800 [6.800, 12.700] | 8.700 [6.700, 12.400] | 9.200 [7.000, 14.500] | 0.22 |
| Carcinoembryonic antigen, ng/mL | – | 1.960 [1.320, 2.850] | 1.830 [1.250, 2.680] | 2.390 [1.600, 4.180] | <0.001*** |
Continuous variables are presented as median [IQR]; categorical variables are expressed as frequency (percentage). ***, P<0.001. CA153, cancer antigen 153; HER2, human epidermal growth factor receptor 2; IQR, interquartile range.
Table 2
| Characteristics | Sort | All | HER2-zero | HER2-low | P |
|---|---|---|---|---|---|
| Gland composition | 25–50% | 211 (50.478) | 116 (53.704) | 95 (47.030) | 0.22 |
| 51–75% | 156 (37.321) | 76 (35.185) | 80 (39.604) | ||
| <25% | 45 (10.766) | 23 (10.648) | 22 (10.891) | ||
| >75% | 6 (1.435) | 1 (0.463) | 5 (2.475) | ||
| Calcification | Contain | 296 (70.813) | 153 (70.833) | 143 (70.792) | 0.99 |
| Not contain | 122 (29.187) | 63 (29.167) | 59 (29.208) | ||
| Age, years | – | 53.000 [46.000, 61.000] | 55.000 [47.000, 61.000] | 52.000 [46.000, 59.000] | 0.08 |
| Maximum diameter, cm | – | 2.200 [1.700, 2.800] | 2.000 [1.500, 2.500] | 2.500 [2.000, 3.500] | <0.001*** |
| CA153, U/mL | – | 8.700 [6.700, 12.400] | 8.700 [6.700, 12.600] | 8.700 [6.800, 11.800] | 0.92 |
| Carcinoembryonic antigen, ng/mL | – | 1.830 [1.250, 2.680] | 1.760 [1.180, 2.420] | 1.930 [1.290, 3.000] | 0.007** |
Continuous variables are presented as median [IQR]; categorical variables are expressed as frequency (percentage). **, P<0.01; ***, P<0.001. CA153, cancer antigen 153; HER2, human epidermal growth factor receptor 2; IQR, interquartile range.
There were 284 cases in the Nomogram (A) training set, 122 cases in the internal validation set, and 137 cases in the external validation set. There were 215 cases in the Nomogram (a) training set, 93 cases in the internal validation set and 110 cases in the external validation set. Associations between HER2 expression and patients’ clinical characteristics as well as mammography features were evaluated. HER2 expression was significantly associated with maximum tumor diameter and CEA (P<0.05).
Feature extraction and derivation of the Rad-score
The ResNet-50 model pre-trained on ImageNet was used to extract DL features. The MLO and CC views were used to extract 2,048 DL features for each view. Subsequently, PCA was applied to reduce feature dimensionality, compressing the features into 32 principal components for each view, which were then concatenated for subsequent analysis. The Pearson correlation coefficient was calculated to assess correlations between features, and for any pair of features with a correlation coefficient greater than 0.9, one redundant feature was removed. The remaining DL features were further selected using the LASSO regression with 10-fold cross-validation. The final models retained 5 and 15 DL features, respectively.
The Rad-score was calculated as a linear combination of the features selected by LASSO (Figure 2):
Uni- and multivariate analyses
In the model training cohort for discriminating between HER2-negative and HER2-positive tumors, univariable analysis showed that maximum tumor diameter and CEA level were significantly associated with HER2 status (P<0.05). Furthermore, the proportion of certain glandular composition categories was 0.00, resulting in extremely wide CIs and indicating potential data sparsity or quasi-complete separation. To mitigate model instability, this variable was excluded from the final multivariable model. In multivariable analysis, both maximum tumor diameter [odds ratio (OR) =1.59, 95% CI: 1.27–1.99, P<0.001], CEA (OR =1.33, 95% CI: 1.12–1.58, P=0.001) were identified as independent predictors of HER2 status (Table 3).
Table 3
| Variables | Univariate analysis | Multivariate analysis | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| β | S.E | Z | P | OR (95% CI) | β | S.E | Z | P | OR (95% CI) | ||
| Gland composition | |||||||||||
| <25% | Reference | ||||||||||
| >75% | −13.18 | 624.19 | −0.02 | 0.99 | 0.00 (0.00–Inf) | ||||||
| 25–50% | 0.68 | 0.46 | 1.50 | 0.13 | 1.98 (0.81–4.84) | ||||||
| 51–75% | 0.03 | 0.49 | 0.05 | 0.96 | 1.03 (0.39–2.71) | ||||||
| Calcification | |||||||||||
| Not contain | Reference | ||||||||||
| Contain | 0.42 | 0.30 | 1.40 | 0.16 | 1.52 (0.85–2.73) | ||||||
| Age | 0.01 | 0.01 | 0.72 | 0.47 | 1.01 (0.98–1.03) | ||||||
| Maximum diameter | 0.51 | 0.11 | 4.54 | <0.001*** | 1.67 (1.34–2.08) | 0.47 | 0.11 | 4.07 | <0.001*** | 1.59 (1.27–1.99) | |
| CA153 | 0.03 | 0.02 | 1.08 | 0.28 | 1.03 (0.98–1.08) | ||||||
| Carcinoembryonic antigen | 0.31 | 0.09 | 3.63 | <0.001*** | 1.36 (1.15–1.61) | 0.28 | 0.09 | 3.18 | 0.001** | 1.33 (1.12–1.58) | |
**, P<0.01; ***, P<0.001. CA153, cancer antigen 153; CI, confidence interval; HER2, human epidermal growth factor receptor 2; Inf, infinity; OR, odds ratio; S.E, standard error.
In the model training cohort for discriminating between HER2-negative and HER2-positive tumors, univariable analysis showed that maximum tumor diameter and CEA level were significantly associated with HER2 status (P<0.05). In multivariable analysis, the maximum tumor diameter (OR =2.44, 95% CI: 1.73–3.43, P<0.001), CEA (OR =1.43, 95% CI: 1.12–1.82, P=0.004) independently predicted two subtypes of HER2-negative (Table 4).
Table 4
| Variables | Univariate analysis | Multivariate analysis | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| β | S.E | Z | P | OR (95% CI) | β | S.E | Z | P | OR (95% CI) | ||
| Gland composition | |||||||||||
| <25% | Reference | ||||||||||
| >75% | 0.33 | 1.30 | 0.25 | 0.80 | 1.38 (0.11–17.67) | ||||||
| 25–50% | −0.56 | 0.48 | −1.18 | 0.24 | 0.57 (0.22–1.45) | ||||||
| 51–75% | −0.18 | 0.48 | −0.37 | 0.71 | 0.83 (0.32–2.16) | ||||||
| Calcification | |||||||||||
| Contain | Reference | ||||||||||
| Not contain | −0.25 | 0.29 | −0.87 | 0.38 | 0.78 (0.44–1.38) | ||||||
| Age | −0.02 | 0.01 | −1.28 | 0.20 | 0.98 (0.96–1.01) | ||||||
| Maximum diameter | 0.88 | 0.17 | 5.11 | <0.001*** | 2.41 (1.72–3.39) | 0.89 | 0.17 | 5.12 | <0.001*** | 2.44 (1.73–3.43) | |
| CA153 | −0.02 | 0.03 | −0.85 | 0.40 | 0.98 (0.92–1.03) | ||||||
| Carcinoembryonic antigen | 0.32 | 0.12 | 2.67 | 0.008** | 1.37 (1.09–1.73) | 0.36 | 0.12 | 2.86 | 0.004** | 1.43 (1.12–1.82) | |
**, P<0.01; ***, P<0.001. CA153, cancer antigen 153; CI, confidence interval; HER2, human epidermal growth factor receptor 2; Inf, infinity; OR, odds ratio; S.E, standard error.
Nomogram construction and performance
A nomogram combining Rad-score and other independent predictors was established to predict HER2 status. Figure 3 lists all three models’ performance indicators, and the predictive performance of them is summarized in Tables 5,6.
Table 5
| Model | AUC (95% CI) | Sensitivity | Specificity | Youden index |
|---|---|---|---|---|
| Training set | ||||
| Clinical model | 0.725 (0.659–0.798) | 0.571 | 0.802 | 0.373 |
| Mammography model | 0.912 (0.864–0.941) | 0.909 | 0.792 | 0.701 |
| Nomogram | 0.934 (0.904–0.964) | 0.922 | 0.802 | 0.724 |
| Internal validation set | ||||
| Clinical model | 0.682 (0.552–0.791) | 0.857 | 0.495 | 0.352 |
| Mammography model | 0.803 (0.652–0.905) | 0.762 | 0.832 | 0.594 |
| Nomogram | 0.913 (0.852–0.965) | 0.905 | 0.832 | 0.736 |
| External validation set | ||||
| Clinical model | 0.719 (0.619–0.819) | 0.667 | 0.691 | 0.358 |
| Mammography model | 0.826 (0.725–0.910) | 0.852 | 0.673 | 0.525 |
| Nomogram | 0.875 (0.792–0.958) | 0.889 | 0.745 | 0.634 |
AUC, area under the curve; CI, confidence interval; HER2, human epidermal growth factor receptor 2.
Table 6
| Model | AUC (95% CI) | Sensitivity | Specificity | Youden index |
|---|---|---|---|---|
| Training set | ||||
| Clinical model | 0.739 (0.663–0.801) | 0.468 | 0.887 | 0.355 |
| Mammography model | 0.948 (0.915–0.969) | 0.945 | 0.858 | 0.803 |
| Nomogram | 0.962 (0.936–0.979) | 0.963 | 0.849 | 0.812 |
| Internal validation set | ||||
| Clinical model | 0.711 (0.611–0.822) | 0.818 | 0.551 | 0.369 |
| Mammography model | 0.849 (0.785–0.917) | 0.773 | 0.776 | 0.548 |
| Nomogram | 0.905 (0.846–0.955) | 0.886 | 0.755 | 0.641 |
| External validation set | ||||
| Clinical model | 0.711 (0.610–0.800) | 0.510 | 0.869 | 0.379 |
| Mammography model | 0.801 (0.694–0.878) | 0.837 | 0.738 | 0.574 |
| Nomogram | 0.882 (0.795–0.936) | 0.776 | 0.902 | 0.677 |
AUC, area under the curve; CI, confidence interval; HER2, human epidermal growth factor receptor 2.
The predictive performance of the three models for distinguishing between HER2-negative and HER2-positive tumors is presented in Figure 3 and Table 5. ROC curves were generated for the training, internal validation, and external validation cohorts to evaluate model performance. Among the three models, the nomogram demonstrated the highest diagnostic efficiency, achieving AUCs of 0.913 and 0.875 in the internal and external validation cohorts, respectively, outperforming both the clinical and DL models. Furthermore, the nomogram had better specificity and sensitivity. The DL model achieved AUCs of 0.803 and 0.826 in the internal and external validation cohorts, respectively, which were higher than those of the clinical model (AUCs of 0.682 and 0.719). Pairwise comparisons of AUCs using the DeLong test are presented in Table 7. The DeLong test indicated that the nomogram achieved significantly higher AUCs than the other two models across all cohorts (all P<0.05). DCA further demonstrated that the nomogram outperformed clinical and DL models, especially at the optimal threshold probability.
Table 7
| Model | Clinical model | Mammography model | Nomogram |
|---|---|---|---|
| Training set | |||
| Clinical model | <0.001 | <0.001 | |
| Mammography model | <0.001 | 0.03 | |
| Nomogram | <0.001 | 0.03 | |
| Internal validation set | |||
| Clinical model | 0.25 | 0.001 | |
| Mammography model | 0.25 | 0.03 | |
| Nomogram | 0.001 | 0.03 | |
| External validation set | |||
| Clinical model | 0.15 | 0.009 | |
| Mammography model | 0.15 | 0.01 | |
| Nomogram | 0.009 | 0.01 |
The values represent P values for pairwise comparisons of AUCs. AUC, area under the curve; HER2, human epidermal growth factor receptor 2.
The predictive performance of the three models in distinguishing HER2-zero from HER2-low tumors is presented in Figure 3 and Table 6. ROC curves for the three models were shown for the training, internal validation, and external validation groups. The nomogram model again achieved the highest diagnostic efficiency, with AUCs of 0.905 and 0.882 in the internal and external validation cohorts, respectively, exceeding the performance of both the clinical and DL models. The DL model achieved AUCs of 0.849 and 0.801 in the internal and external validation cohorts, while the clinical model showed an AUC of 0.711 in both cohorts. Pairwise comparisons of AUCs using the DeLong test are presented in Table 8, demonstrating that the nomogram significantly outperformed the other two models across cohorts (all P<0.05). DCA further confirmed that the nomogram outperformed clinical and DL models, especially at the optimal threshold probability. Calibration curves showed good agreement between predicted probabilities and observed outcomes, indicating robust model calibration.
Table 8
| Model | Clinical model | Mammography model | Nomogram |
|---|---|---|---|
| Training set | |||
| Clinical model | <0.001 | <0.001 | |
| Mammography model | <0.001 | 0.048 | |
| Nomogram | <0.001 | 0.048 | |
| Internal validation set | |||
| Clinical model | 0.057 | 0.001 | |
| Mammography model | 0.057 | 0.002 | |
| Nomogram | 0.001 | 0.002 | |
| External validation set | |||
| Clinical model | 0.23 | 0.002 | |
| Mammography model | 0.23 | 0.007 | |
| Nomogram | 0.002 | 0.007 |
The values represent P values for pairwise comparisons of AUCs. AUC, area under the curve; HER2, human epidermal growth factor receptor 2.
Discussion
HER2 status has traditionally been regarded as a binary classification (positive or negative) and has guided the use of anti-HER2-directed therapies, most commonly trastuzumab. However, the advent of ADCs in recent years has fundamentally reshaped the clinical interpretation of HER2 expression. New-generation ADCs, exemplified by T-DXd, have demonstrated significant improvements in PFS and OS compared with conventional chemotherapy in previously treated patients with metastatic breast cancer and HER2-low expression (IHC 1+ or IHC 2+/ISH−), as shown in the pivotal phase III DESTINY-Breast04 trial (17,25). These landmark findings directly led to the approval of T-DXd for the treatment of HER2-low breast cancer by the US Food and Drug Administration and other regulatory agencies worldwide, thereby formally establishing HER2-low disease as an independent biological subtype with clear therapeutic relevance (14). Consequently, accurate discrimination between HER2-low and HER2-zero tumors has become critical for identifying patients who may benefit from these transformative therapies. Tumor clinicopathological characteristics play a central role in guiding optimal treatment strategies. DL-based imaging analysis enables rapid, non-invasive assessment of the entire tumor burden and associated lesions. Therefore, this retrospective study aimed to investigate whether DL models derived from mammography could be used to characterize the pathological features of malignant breast lesions. Previous studies combining mammography and DL have predominantly focused on the binary classification of HER2 status (negative versus positive) in breast cancer. Luo et al. (21) proposed a DenseNet-121 model integrated with a convolutional block attention module (CBAM) to predict multiple molecular subtypes of breast cancer, including binary HER2 status, based on mammography. In an independent test set, the AUC for HER2 prediction was 0.658, and Grad-CAM visualization identified the peritumoral region as a key discriminative feature. Zeng et al. (22) developed a CBAM-enhanced ResNet-18 model to simultaneously predict the expression of HER2 (binary classification), estrogen receptor (ER), and progesterone receptor (PR) using whole mammographic images without manual tumor segmentation, achieving an AUC of 0.708 for HER2 prediction. Qiu et al. (23) proposed a framework for binary HER2 expression prediction by integrating multi-sequence MRI (T1-weighted, T2-weighted, and dynamic contrast-enhanced images) with DL models, including ResNet-50 and VGG16, and subsequently constructed a nomogram using intraclass correlation coefficient filtering and LASSO feature selection. This approach achieved an AUC of 0.94 based on a multicenter dataset comprising 6,438 cases. Collectively, these studies provided important evidence supporting the feasibility of imaging-based DL approaches for HER2-positive versus HER2-negative classification and contributed to the optimization of targeted therapy strategies in breast cancer. However, none of these studies specifically addressed the clinically relevant sub-differentiation between HER2-low expression and HER2-zero expression (IHC 0), which has emerged as a key determinant for ADC-based treatment selection.
CEA is one of the most commonly used serum biomarkers in breast cancer, and elevated serum CEA levels are frequently observed in affected patients. CEA is considered a reliable indicator for monitoring distant metastasis, disease recurrence after treatment, and therapeutic response evaluation (26,27). Biologically, CEA has been shown to increase the growth and survival of cancer cells. Overexpression of CEA can facilitate the detachment of tumor cells from surrounding stromal cells and normal cells, diminish the cellular and extracellular matrix interactions, and decrease their adhesion, thereby enhancing tumor invasiveness. Previous studies have demonstrated that serum CEA levels are positively correlated with the TNM stage, histological grade, and lymph node metastasis in breast cancer (28,29). Tumor size is directly reflected in the T stage of the TNM classification system. Specifically, T1 tumors have a maximum diameter of ≤2 cm, T2 tumors measure >2 cm but ≤5 cm, and T3 tumors exceed 5 cm in maximum diameter. As tumor size increases, disease stage generally advances, which is associated with a poorer prognosis. The observed association between maximum tumor diameter and HER2 positivity may reflect the aggressive biological behavior of HER2-enriched tumors, which are typically characterized by rapid growth and heightened proliferative activity (30,31). This study identified tumor maximum diameter and CEA as independent risk factors that can be used to predict the expression status of the HER2 gene. When incorporated into the clinical model, these parameters yielded AUC values of 0.682 and 0.719 for distinguishing HER2-negative from HER2-positive tumors in the internal and external validation cohorts, respectively, and 0.711 and 0.711 for distinguishing HER2-zero from HER2-low tumors, indicating modest predictive performance. In univariate logistic regression analyses of the two training cohorts, glandular composition, calcification, age, and CA15-3 were assessed but did not reach statistical significance (P>0.05) and were therefore excluded from the final model. Collectively, these findings suggest that, compared with maximum tumor diameter and CEA, these variables provide limited independent predictive value for HER2 status.
In contrast to several of the aforementioned studies that were limited to DL models alone, this multicenter study further integrated DL-derived features with key clinical variables, including maximum tumor diameter and CEA, to construct two visual nomograms for predicting HER2 status in patients with breast cancer. Notably, HER2-low and HER2-zero expression were further differentiated and independently predicted. This strategy translates complex artificial intelligence outputs into intuitive risk assessment tools that are readily interpretable and applicable in routine clinical practice, thereby facilitating model interpretability, translational potential, and clinical adoption. The nomogram models consistently outperformed models based solely on clinical variables or DL features in predicting HER2 status. Moreover, external validation using independent datasets further confirmed the robustness, generalizability, and clinical applicability of the proposed models. Collectively, these findings suggest that the developed nomograms may serve as non-invasive adjunctive tools for the early assessment of HER2 status prior to pathological confirmation, thereby providing valuable decision support for individualized treatment selection.
Unlike previous studies that relied on manual ROI delineation, the present study utilized whole mammographic images, thereby improving clinical feasibility and reducing interobserver variability. DL is capable of effectively extracting a large number of quantitative imaging features that indicate characteristics such as texture, intensity, heterogeneity, and morphology (32). Although these features are not visually discernible, they can capture intratumoral heterogeneity at the cellular level. By analyzing the microstructural changes in the tumor region, DL can accurately predict HER2 status. Currently, a variety of machine learning algorithms have been employed for radiomics analysis. LightGBM is a highly efficient framework for gradient boosting, as described in previous studies (33). The design of this system incorporates characteristics such as fast processing speed, minimal memory usage, exceptional accuracy, and the ability to perform parallel computing, enabling it to effectively handle extensive amounts of data. The algorithm of this system is based on histograms, and it uses a growth strategy called “leaf-wise” and supports sparse features. These features make it exceptional in machine learning competitions and in practical applications in industries. The AUC of the LightGBM model exceeded 0.8 in both the training cohort and the internal and external validation cohort in our study. These findings suggest that the proposed model shows promising capability for predicting HER2 status in breast cancer patients and may contribute to the development of individualized treatment strategies in clinical practice.
With the demonstrated therapeutic efficacy of ADCs in patients with HER2-low breast cancer, a primary objective of this study was to develop a model capable of discriminating between HER2-low and HER2-zero expression, a distinction with substantial and increasingly important clinical implications. Nevertheless, predictive models specifically addressing HER2-low and HER2-zero breast cancer remain limited. Accurate pre-treatment identification of absent or minimal HER2 expression is therefore critical for optimizing treatment stratification and guiding therapeutic decision-making. In this study, the DL model demonstrated moderate performance in predicting the three HER2 expression categories. By integrating the Rad-score with clinical variables, including CEA levels and maximum tumor diameter, a nomogram was constructed, resulting in further improvement in predictive performance, as reflected by increased AUC values. Moreover, validation using independent external datasets supported the robustness, generalizability, and potential applicability of the model in real-world clinical settings. DCA demonstrated that the nomogram provided a greater net clinical benefit in predicting the expression status of HER2 before surgery compared to other models in the training cohort, internal validation cohort, and external validation cohort. Meanwhile, DeLong test results showed that the AUC of the nomogram was significantly higher than that of each unimodal model across all cohorts (all P<0.05), indicating superior discriminative performance. However, the analysis primarily focused on pairwise comparisons between the nomogram and individual unimodal models. Given the exploratory nature of this study and the limited number of comparisons, no adjustment for multiple testing was applied.
Accurate preoperative prediction of HER2 status has important implications for surgical planning and therapeutic decision-making. Previous studies have demonstrated that preoperative assessment of HER2 expression is closely associated with treatment selection and prognosis in breast cancer, and that its integration with other clinical indicators can facilitate the development of individualized treatment strategies (34). Moreover, neoadjuvant therapy—particularly regimens incorporating HER2-targeted agents—has been shown to substantially reduce tumor burden in patients with HER2-positive breast cancer, increase the likelihood of breast-conserving surgery, and potentially decrease the need for conventional axillary surgical procedures. Collectively, these findings highlight the clinical relevance of preoperative HER2 assessment for risk stratification, treatment optimization, and surgical planning (35).
This multicenter study has several inherent limitations that should be acknowledged. First, subgroup analyses based on tumor stage and HR status were not performed. Future studies with larger cohorts are needed to evaluate the predictive performance of the model across diverse clinical subgroups. In addition, the exclusion of IHC 2+ cases not confirmed by FISH, together with the requirement for complete imaging and pathological data, may introduce selection bias and limit the generalizability of the model to real-world clinical populations. Prospective validation in broader, more heterogeneous cohorts is therefore warranted. Second, some variables exhibited data sparsity in the univariate analysis, underscoring the need for verification in larger samples. Third, the ResNet-50 model was pre-trained on natural images (ImageNet) without domain-specific fine-tuning for mammography. Although this may introduce subtle domain discrepancies between extracted features and breast lesion representations, this approach was chosen to reduce the risk of overfitting on the limited training data while leveraging shared visual features such as texture and edge information. Future studies may improve feature representation and predictive performance by pre-training on large-scale breast imaging datasets or applying domain-adaptive fine-tuning strategies. To overcome the limitation of sample size, we plan to expand the dataset by collecting additional samples from multiple medical institutions. Furthermore, integrating multimodal imaging data with genomic sequencing information is expected to increase feature dimensionality and further enhance model performance.
To summarize, we developed two binary nomograms integrating DL-derived imaging features with clinical variables to predict three HER2 expression subtypes in patients with breast cancer. This approach enables non-invasive, pre-treatment assessment of HER2 status prior to pathological confirmation, highlighting the potential of artificial intelligence to support personalized treatment decision-making in breast cancer.
Conclusions
In this study, we developed two nomograms integrating DL-derived imaging features with clinical variables to predict HER2 status. Notably, the models could distinguish not only between HER2-positive and HER2-negative tumors but also between HER2-low and HER2-zero subtypes. Compared with unimodal models, the nomograms demonstrated superior predictive performance and good generalizability in external validation. Clinically, they may serve as non-invasive tools for preoperative HER2 assessment, supporting treatment planning and patient stratification. In particular, identifying HER2-low tumors may help select patients who could benefit from ADC therapy. With further prospective validation, these models could be incorporated into routine clinical workflows.
Acknowledgments
None.
Footnote
Reporting Checklist: The authors have completed the TRIPOD reporting checklist. Available at https://gs.amegroups.com/article/view/10.21037/gs-2025-aw-467/rc
Data Sharing Statement: Available at https://gs.amegroups.com/article/view/10.21037/gs-2025-aw-467/dss
Peer Review File: Available at https://gs.amegroups.com/article/view/10.21037/gs-2025-aw-467/prf
Funding: This work was supported by
Conflicts of Interest: All authors have completed the ICMJE uniform disclosure form (available at https://gs.amegroups.com/article/view/10.21037/gs-2025-aw-467/coif). The authors have no conflicts of interest to declare.
Ethical Statement: The authors are accountable for all aspects of the work in ensuring that questions related to the accuracy or integrity of any part of the work are appropriately investigated and resolved. The study was conducted in accordance with the Declaration of Helsinki and its subsequent amendments. The study was approved by the Ethics Committee of Affiliated Hospital of Jining Medical University (ethics approval No. 2021C019) and individual consent for this retrospective analysis was waived.
Open Access Statement: This is an Open Access article distributed in accordance with the Creative Commons Attribution-NonCommercial-NoDerivs 4.0 International License (CC BY-NC-ND 4.0), which permits the non-commercial replication and distribution of the article with the strict proviso that no changes or edits are made and the original work is properly cited (including links to both the formal publication through the relevant DOI and the license). See: https://creativecommons.org/licenses/by-nc-nd/4.0/.
References
- Lamb CA, Vanzulli SI, Lanari C. Hormone receptors in breast cancer: more than estrogen receptors. Medicina (B Aires) 2019;79:540-5.
- Yarden Y. Biology of HER2 and its importance in breast cancer. Oncology 2001;61:1-13. [Crossref] [PubMed]
- Banerji U, van Herpen CML, Saura C, et al. Trastuzumab duocarmazine in locally advanced and metastatic solid tumours and HER2-expressing breast cancer: a phase 1 dose-escalation and dose-expansion study. Lancet Oncol 2019;20:1124-35. [Crossref] [PubMed]
- Wolff AC, Hammond MEH, Allison KH, et al. Human Epidermal Growth Factor Receptor 2 Testing in Breast Cancer: American Society of Clinical Oncology/College of American Pathologists Clinical Practice Guideline Focused Update. J Clin Oncol 2018;36:2105-22. [Crossref] [PubMed]
- Hamilton E, Shastry M, Shiller SM, et al. Targeting HER2 heterogeneity in breast cancer. Cancer Treat Rev 2021;100:102286. [Crossref] [PubMed]
- Yang J, Ju J, Guo L, et al. Prediction of HER2-positive breast cancer recurrence and metastasis risk from histopathological images and clinical information via multimodal deep learning. Comput Struct Biotechnol J 2022;20:333-42. [Crossref] [PubMed]
- Shui R, Liang X, Li X, et al. Hormone Receptor and Human Epidermal Growth Factor Receptor 2 Detection in Invasive Breast Carcinoma: A Retrospective Study of 12,467 Patients From 19 Chinese Representative Clinical Centers. Clin Breast Cancer 2020;20:e65-74. [Crossref] [PubMed]
- Ishii K, Morii N, Yamashiro H. Pertuzumab in the treatment of HER2-positive breast cancer: an evidence-based review of its safety, efficacy, and place in therapy. Core Evid 2019;14:51-70. [Crossref] [PubMed]
- Takada M, Toi M. Neoadjuvant treatment for HER2-positive breast cancer. Chin Clin Oncol 2020;9:32. [Crossref] [PubMed]
- Geyer CE, Forster J, Lindquist D, et al. Lapatinib plus capecitabine for HER2-positive advanced breast cancer. N Engl J Med 2006;355:2733-43. [Crossref] [PubMed]
- Yuan Y, Liu X, Cai Y, et al. Lapatinib and lapatinib plus trastuzumab therapy versus trastuzumab therapy for HER2 positive breast cancer patients: an updated systematic review and meta-analysis. Syst Rev 2022;11:264. [Crossref] [PubMed]
- Wang X, Wang L, Yu Q, et al. The Effectiveness of Lapatinib in HER2-Positive Metastatic Breast Cancer Patients Pretreated With Multiline Anti-HER2 Treatment: A Retrospective Study in China. Technol Cancer Res Treat 2021;20:15330338211037812. [Crossref] [PubMed]
- Dieci MV, Miglietta F. HER2: a never ending story. Lancet Oncol 2021;22:1051-2. [Crossref] [PubMed]
- Tarantino P, Hamilton E, Tolaney SM, et al. HER2-Low Breast Cancer: Pathological and Clinical Landscape. J Clin Oncol 2020;38:1951-62. [Crossref] [PubMed]
- Ferraro E, Drago JZ, Modi S. Implementing antibody-drug conjugates (ADCs) in HER2-positive breast cancer: state of the art and future directions. Breast Cancer Res 2021;23:84. [Crossref] [PubMed]
- Najjar MK, Manore SG, Regua AT, et al. Antibody-Drug Conjugates for the Treatment of HER2-Positive Breast Cancer. Genes (Basel) 2022;13:2065. [Crossref] [PubMed]
- Modi S, Jacot W, Iwata H, et al. Trastuzumab deruxtecan in HER2-low metastatic breast cancer: long-term survival analysis of the randomized, phase 3 DESTINY-Breast04 trial. Nat Med 2025;31:4205-13. [Crossref] [PubMed]
- Bagegni NA, Giridhar KV, Stewart D. HER2-Low and HER2-Ultralow Metastatic Breast Cancer and Trastuzumab Deruxtecan: Common Clinical Questions and Answers. Cancers (Basel) 2025;17:4021. [Crossref] [PubMed]
- Zhou J, Tan H, Li W, et al. Radiomics Signatures Based on Multiparametric MRI for the Preoperative Prediction of the HER2 Status of Patients with Breast Cancer. Acad Radiol 2021;28:1352-60. [Crossref] [PubMed]
- Xu A, Chu X, Zhang S, et al. Development and validation of a clinicoradiomic nomogram to assess the HER2 status of patients with invasive ductal carcinoma. BMC Cancer 2022;22:872. [Crossref] [PubMed]
- Luo Y, Wei J, Gu Y, et al. Predicting molecular subtype in breast cancer using deep learning on mammography images. Front Oncol 2025;15:1638212. [Crossref] [PubMed]
- Zeng S, Chen H, Jing R, et al. An assessment of breast cancer HER2, ER, and PR expressions based on mammography using deep learning with convolutional neural networks. Sci Rep 2025;15:4826. [Crossref] [PubMed]
- Qiu S, Zhao Q, Zhao Y. Novel deep learning-based prediction of HER2 expression in breast cancer using multimodal MRI, nomogram, and decision curve analysis. Front Oncol 2025;15:1593033. [Crossref] [PubMed]
- Wolff AC, Somerfield MR, Dowsett M, et al. Human Epidermal Growth Factor Receptor 2 Testing in Breast Cancer: ASCO-College of American Pathologists Guideline Update. J Clin Oncol 2023;41:3867-72. [Crossref] [PubMed]
- Modi S, Jacot W, Yamashita T, et al. Trastuzumab Deruxtecan in Previously Treated HER2-Low Advanced Breast Cancer. N Engl J Med 2022;387:9-20. [Crossref] [PubMed]
- Tarighati E, Keivan H, Mahani H. A review of prognostic and predictive biomarkers in breast cancer. Clin Exp Med 2023;23:1-16. [Crossref] [PubMed]
- Fan Y, Chen X, Li H. Clinical value of serum biomarkers CA153, CEA, and white blood cells in predicting sentinel lymph node metastasis of breast cancer. Int J Clin Exp Pathol 2020;13:2889-94.
- Li X, Dai D, Chen B, et al. Clinicopathological and Prognostic Significance of Cancer Antigen 15-3 and Carcinoembryonic Antigen in Breast Cancer: A Meta-Analysis including 12,993 Patients. Dis Markers 2018;2018:9863092. [Crossref] [PubMed]
- Seale KN, Tkaczuk KHR. Circulating Biomarkers in Breast Cancer. Clin Breast Cancer 2022;22:e319-31. [Crossref] [PubMed]
- Omranipour R, Nazarian N, Alipour S, et al. Evaluation of HER2 Positivity Based on Clinicopathological Findings in HER2 Borderline Tumors in Iranian Patients with Breast Cancer. Iran J Pathol 2023;18:403-9. [Crossref] [PubMed]
- Cheng X. A Comprehensive Review of HER2 in Cancer Biology and Therapeutics. Genes (Basel) 2024;15:903. [Crossref] [PubMed]
- Yip SS, Aerts HJ. Applications and limitations of radiomics. Phys Med Biol 2016;61:R150-66. [Crossref] [PubMed]
- Liao H, Zhang X, Zhao C, et al. LightGBM: an efficient and accurate method for predicting pregnancy diseases. J Obstet Gynaecol 2022;42:620-9. [Crossref] [PubMed]
- Zhao Z, Yuan H, Song X, et al. Preoperative prediction of HER2 expression and sentinel lymph node status in breast cancer using a mammography radiomics model. Front Oncol 2025;15:1578458. [Crossref] [PubMed]
- NICE Evidence Reviews Collection. Neoadjuvant chemotherapy for people with HER2 positive breast cancer: Early and locally advanced breast cancer: diagnosis and management: Evidence review T. London: National Institute for Health and Care Excellence (NICE); 2025.

