Interpretable machine learning for prognostic prediction in young women with triple-negative invasive ductal breast cancer: a Boruta-SHAP integrated approach
Original Article

Interpretable machine learning for prognostic prediction in young women with triple-negative invasive ductal breast cancer: a Boruta-SHAP integrated approach

Lili Luo1, Shulian Li2, Qiang Ji3 ORCID logo

1Department of Anesthesiology, West China Hospital, Sichuan University, Chengdu, China; 2Department of Thyroid and Breast Surgery, West China Hospital, Sichuan University, Chengdu, China; 3Department of Aesthetic Plastic Surgery, West China School of Public Health and West China Fourth Hospital, Sichuan University, Chengdu, China

Contributions: (I) Conception and design: Q Ji; (II) Administrative support: Q Ji; (III) Provision of study materials or patients: L Luo; (IV) Collection and assembly of data: L Luo, S Li; (V) Data analysis and interpretation: L Luo, S Li; (VI) Manuscript writing: All authors; (VII) Final approval of manuscript: All authors.

Correspondence to: Qiang Ji, MD. Department of Aesthetic Plastic Surgery, West China School of Public Health and West China Fourth Hospital, Sichuan University, No. 16, Section 3, Renmin South Road, Wuhou District, Chengdu 610041, China. Email: 2021441662@qq.com.

Background: Young women with triple-negative invasive ductal breast cancer (TN-IDC) exhibit high invasiveness and poor prognosis, rendering accurate prognostic prediction crucial for individualized diagnosis and treatment. This study aimed to identify core prognostic factors and establish an optimal binary predictive model for cancer-specific survival (CSS) through an interpretable machine learning (ML) framework comprising “feature selection-model construction-result interpretation”.

Methods: Clinical data of 2,311 young female TN-IDC patients aged 20–40 years were extracted from the Surveillance, Epidemiology, and End Results (SEER) database, including 560 deceased cases and 1,751 surviving cases. The dataset covered multi-dimensional indicators such as demographics, tumor characteristics, treatment regimens, and metastatic status. The Boruta algorithm was used to screen key prognostic features for CSS, and 8 ML models were established based on the selected features. Model performance was comprehensively evaluated using 9 metrics (including accuracy, sensitivity, and specificity) as well as receiver operating characteristic (ROC) curves and calibration curves. Finally, the Shapley Additive Explanations (SHAP) algorithm was employed to interpret model logic and quantify feature contributions to CSS.

Results: Boruta feature selection identified node (N) stage, tumor (T) stage, bone, lung, and liver metastases, and surgical approach as core prognostic factors for CSS. Among the 8 established ML models, the multilayer perceptron (MLP) model demonstrated the optimal overall performance, with an accuracy of 0.711 [95% confidence interval (CI): 0.683–0.739], Matthews Correlation Coefficient of 0.402 (95% CI: 0.372–0.423), balanced accuracy of 0.732 (95% CI: 0.704–0.759), area under the curve (AUC) of 0.782 (95% CI: 0.737–0.828). Additionally, its calibration curve showed the highest alignment with the ideal line. Several models achieved competitive performance, including logistic regression (AUC =0.736), xgboost (AUC =0.741), random forest (rf, AUC =0.726) and elastic net (enet, AUC =0.738), yet none surpassed MLP in overall predictive ability. AUC differences between MLP and logistic regression (P=0.17) or xgboost (P=0.22) were not statistically significant, but MLP showed consistently superior performance across multiple metrics, confirming its robustness for this task. SHAP analysis revealed that the N3 stage (N stage subclass) and bone metastasis had the most significant impacts on predictive outcomes, with higher N stages and positive bone metastasis status significantly increasing the risk of adverse prognosis.

Conclusions: Through the Boruta-SHAP interpretable ML framework, this study clarified the core prognostic characteristics of TN-IDC in young women. The constructed MLP model shows favorable prognostic performance, providing supportive evidence and a supplementary practical tool for clinical risk stratification, individualized treatment decision-making, and follow-up management.

Keywords: Triple-negative invasive ductal breast cancer (TN-IDC); machine learning (ML); Boruta-Shapley Additive Explanations algorithm (Boruta-SHAP algorithm); multilayer perceptron neural network (MLP neural network)


Submitted Feb 03, 2026. Accepted for publication Apr 15, 2026. Published online May 27, 2026.

doi: 10.21037/gs-2026-1-0091


Highlight box

Key findings

• The Boruta‑Shapley Additive Explanations (SHAP) interpretable machine learning framework was used to develop prognostic models for young women with triple‑negative invasive ductal breast cancer (TN‑IDC). The multilayer perceptron (MLP) model showed optimal predictive performance and good calibration. Node (N) stage, tumor (T) stage, bone/lung/liver metastases, and surgical approach were identified as core prognostic factors, with N3 stage and bone metastasis having the strongest adverse effects.

What is known and what is new?

• Young TN‑IDC patients have a poor prognosis and lack dedicated, interpretable prognostic tools.

• This study built 8 models using a large Surveillance, Epidemiology, and End Results database and provided transparent interpretation via Boruta‑SHAP, offering a new prediction tool for this high‑risk group.

What is the implication, and what should change now?

• The MLP model may offer supportive information for risk stratification and assist clinical decision‑making for young patients with TN‑IDC. Integration of this interpretable model might help improve clinical management.


Introduction

Breast cancer is the most commonly diagnosed cancer among women worldwide, with substantial increases in new cases and deaths from 1990 to 2023. Although agestandardized mortality rates have declined slightly, breast cancer remains a leading threat to women’s health, with a continuing and growing burden expected through 2050, particularly in low- and middle-income settings (1). Among the numerous subtypes of breast cancer, triple-negative invasive ductal breast cancer (TN-IDC) is particularly challenging in clinical diagnosis and treatment due to the negative expression of estrogen receptor (ER), progesterone receptor (PR), and human epidermal growth factor receptor 2 (HER2), which results in a lack of well-defined targeted therapeutic targets (2). In contrast to other breast cancer subtypes, TN-IDC exhibits distinct biological characteristics of local recurrence, including extremely strong invasiveness, early recurrence, and high metastatic risk: the median time to recurrence after surgery is only 2–3 years, the 5-year local recurrence rate reaches 8–15%, the risk of distant metastasis within 2 years is as high as 60%, and the 5-year overall survival rate is merely 30–45%, indicating an extremely poor prognosis (3).

Notably, significant differences exist in the clinical characteristics of TN-IDC among different age groups. Epidemiological studies have demonstrated that the incidence of TN-IDC is significantly higher in young women (typically defined as ≤40 years old) than in elderly women. TN-IDC accounts for 20–30% of all breast cancer cases in young patients, whereas this proportion decreases to only 15–18% in patients aged ≥60 years (4-6). This phenomenon may be associated with several factors, including the high proliferative activity of breast tissue, genetic susceptibility (e.g., BRCA1/2 gene mutations), hormonal fluctuations, and lifestyle habits in young women (7). More critically, young women with TN-IDC often face worse clinical outcomes. On the one hand, young patients are more likely to be diagnosed at an advanced tumor stage: approximately 30–40% of them already have regional lymph node metastasis or distant metastasis at diagnosis, while the proportion of early-stage diagnosis in elderly patients can exceed 50% (8); on the other hand, young patients develop chemoresistance more rapidly, with a peak of recurrence concentrated within 2–3 years after treatment. Moreover, the preferred sites of metastasis are vital organs such as the lungs, liver, and bone. Once distant metastasis occurs, the available therapeutic options are extremely limited, leading to a significant decline in both quality of life and overall survival (9). Therefore, conducting accurate disease risk assessment and prognosis prediction for young women with TN-IDC has become a crucial breakthrough to improve the survival outcomes of this population.

In traditional clinical practice, physicians primarily rely on indicators such as the American Joint Committee on Cancer (AJCC) tumor staging system and pathological grade to evaluate the prognosis of young women with TN-IDC. However, these methods have obvious limitations: they only focus on core characteristics including tumor size and lymph node metastasis, without fully integrating multi-dimensional information such as treatment regimens and metastatic sites. Additionally, they are highly dependent on subjective clinical experience, resulting in insufficient accuracy of risk stratification after early diagnosis. This not only affects the consistency of individualized treatment decisions but may also lead to overtreatment (e.g., unnecessary chemotherapy for low-risk patients) or undertreatment (e.g., missed intervention opportunities for patients with occult high recurrence risk), thereby increasing the medical burden and physical-mental stress on patients (10-12). Currently, there is a lack of specific prediction tools for this particular population, and the existing models lack validation with large-sample clinical data. Therefore, it is urgently necessary to develop an accurate and objective prediction tool based on easily accessible clinical indicators to optimize risk assessment and diagnosis-treatment pathways after early diagnosis, which can balance clinical efficacy, patient quality of life, and the efficient allocation of medical resources.

In recent years, the application of machine learning (ML) technology in the medical field has provided a new technical approach to address this challenge (13). ML can automatically perform data analysis, feature extraction, and pattern recognition on massive datasets through algorithms. It is particularly adept at processing high-dimensional and nonlinear complex clinical data, thus overcoming the limitations of traditional statistical methods (14). In the field of oncology, ML has been widely applied in various scenarios, including cancer screening (e.g., benign-malignant differentiation of breast imaging), prognosis prediction (e.g., survival risk stratification), and treatment response evaluation (e.g., prediction of chemotherapy sensitivity) (14-16). Nevertheless, existing studies still have certain shortcomings: some models rely on multi-omics data, which are associated with high clinical acquisition costs and great difficulty in popularization; there are few specialized models targeting the special high-risk population of young women; furthermore, most models lack sufficient interpretability analysis, making it difficult to identify the core prognostic driving factors, which limits their clinical translational application. Therefore, constructing a prognosis prediction model for young women with TN-IDC that is based on easily accessible clinical indicators and possesses both high accuracy and good interpretability has important clinical implications for individualized management. We present this article in accordance with the TRIPOD reporting checklist (available at https://gs.amegroups.com/article/view/10.21037/gs-2026-1-0091/rc).


Methods

Study population and data source

This cross-sectional study retrieved data from 18 geographically diverse cancer registries within the Surveillance, Epidemiology, and End Results (SEER) database [National Cancer Institute (NCI)] using SEER*Stat software (Version 8.3.4). All extraction fields and screening procedures strictly adhered to SEER’s official protocols to ensure data scientific validity and compliance. The study was conducted in accordance with the Declaration of Helsinki and its subsequent amendments.

Eligible patients met the following inclusion criteria: (I) female with pathologically confirmed primary invasive ductal breast cancer diagnosed from 2010 to 2020 [International Classification of Diseases for Oncology, 3rd Edition (ICD-O-3)]; the starting year of 2010 was chosen because the SEER database began systematic recording of HER2 status in that year, enabling reliable identification of the triple-negative subtype (17). (II) Aged 20–40 years at diagnosis. (III) Definite triple-negative subtype (ER, PR, HER2). (IV) Complete clinical and follow-up data with cancer-specific survival (CSS) status, enabling extraction of core variables: demographics (race, age, etc.), tumor characteristics [primary site, AJCC 7th Edition tumor-node-metastasis (TNM) stage, etc.], treatment details (surgery, radiotherapy, chemotherapy), and distant metastasis status (bone, brain, liver, lung).

Exclusion criteria were: (I) prior history of other malignant tumors before breast cancer diagnosis; (II) missing tumor staging data; (III) missing core variables (e.g., age, ER/PR/HER2 status) without valid imputation; (IV) ambiguous diagnostic information (e.g., unclear pathological diagnosis) or logical contradictions (e.g., inconsistent staging and tumor size).

Data preprocessing and splitting

Variables were coded standardizedly: multi-categorical variables (e.g., race, primary tumor site, surgical approach) were one-hot encoded to avoid spurious ordinal relationships, while dichotomous variables (e.g., metastasis status, radiotherapy administration) were binary coded (0= absent/not administered, 1= present/administered).

The dataset was split into a training set (n=1,733, 75%) for model building and hyperparameter optimization, and a validation set (n=587, 25%) for independent validation, using stratified randomization (seed =42) to ensure reproducibility. Stratification was based on metastasis and treatment status, maintaining consistent distributions of core variables between sets and reducing bias from class imbalance. Owing to the inherent class imbalance (survivors:deceased ≈3:1), all models were trained with class weight adjustment to mitigate majority-class bias.

Boruta algorithm for feature selection

Key prognostic factors were primarily screened using the Boruta algorithm, which performs feature significance testing by generating biologically meaningless “shadow features” as a baseline reference through random permutation of original feature values (18). A random forest classifier was used as the base model for its stability in estimating feature importance. The two-sided Wilcoxon rank-sum test was applied to compare the importance distribution between original features and shadow features. Features with significance lower than shadow features (P>0.05) were regarded as redundant and removed, retaining only variables with significant predictive value for CSS.

Model construction and model evaluation metrics

Eight ML models were constructed to predict the risk of CSS in young women with TN-IDC. CSS was defined as a binary endpoint: survival (alive from breast cancer) or death (dead due to breast cancer). Models included elastic net (enet), decision tree (dt), random forest (rf), extreme gradient boosting (xgboost), multilayer perceptron (MLP), regularized support vector machine (rsvm), K-nearest neighbors (knn), and logistic regression (logistic).

We used a comprehensive multi-dimensional system to evaluate model performance. Given the class imbalance, balanced accuracy, Matthews correlation coefficient (MCC), and Youden’s index were used as primary metrics instead of overall accuracy to prevent performance overestimation. Key metrics included: (I) accuracy; (II) Cohen’s Kappa (kap); (III) sensitivity (sens); (IV) specificity (spec); (V) positive predictive value (PPV); (VI) negative predictive value (NPV); (VII) MCC; (VIII) Youden’s index (j_index); (IX) balanced accuracy [bal_accuracy, (sens+spec)/2]. Additionally, receiver operating characteristic (ROC) curves [area under the curve (AUC) for discriminative ability] and calibration curves (Brier score for prediction-reality alignment) were generated. The optimal model was determined via holistic comparison of these metrics and curves.

All models were implemented in R software (version 4.4.1). Hyperparameter optimization was carried out using 5-fold cross-validation combined with grid search, and the parameter ranges were set based on previous literature and standard theoretical values. For the MLP model, we tested a set of common settings: hidden layer sizes [16, 24, 32], a standard nonlinear activation function, a widely used adaptive optimization method, learning rates (0.0001, 0.001, 0.01), L2 regularization, and batch sizes [32, 64]. The optimal structure was confirmed as one hidden layer with 24 nodes through 5-fold cross-validation. To handle the class imbalance (75.8% alive, 24.2% deceased), class weights were applied during model training. Synthetic sampling methods such as SMOTE were not used to avoid artificial data that might weaken clinical interpretability and generalization.

Shapley Additive Explanations (SHAP) analysis for model interpretation

SHAP values were used to quantify the contribution of each feature to the prognostic prediction. A larger absolute SHAP value indicates greater feature importance. Positive SHAP values represent an increased risk of poor CSS, while negative values indicate a decreased risk. Based on the absolute SHAP values, all features were ranked to determine their relative importance. This approach enabled transparent and intuitive interpretation of how each clinical feature influenced the final prognosis prediction in young women with triple-negative invasive ductal breast cancer.

Statistical analysis

All statistical analyses were performed using two-tailed tests, with a significance level of α=0.05. A P value <0.01 was considered to indicate a statistically significant difference. Data preprocessing, variable coding, dataset splitting, and descriptive statistical analysis were implemented using R software (Version 4.4.1) combined with Boruta (Version 9.0.0), and fastshap (Version 0.1.1). The calculation of AUC and plotting of ROC curves were performed using the pROC package (Version 1.19.0). The visualization of feature importance and model interpretation was implemented using the NeuralNetTools package (Version 1.5.3) eta.


Results

Clinical and pathological characteristics

A total of 2,311 young female patients with TN-IDC were enrolled in this study in Table 1. Of these, 560 (24.2%) patients died and 1,751 (75.8%) survived. Baseline characteristics between the two groups were compared, including race, tumor location, T/N stage, surgical approach, chemotherapy, radiotherapy, and distant metastasis. White patients constituted the majority (70.4%), while Black patients accounted for a significantly higher proportion in the deceased group (27.3%) than in the survival group (17.1%). Most tumors were located in the upper-outer quadrant (38.6%). The proportion of advanced T stage (T3/T4) and N stage (N2/N3) was significantly higher in the deceased group, while N0 stage was more common in the survival group. Total mastectomy was the most common surgical procedure (39.0%). The deceased group showed higher rates of no surgery and radical mastectomy, whereas the survival group had a higher rate of breast-onserving surgery.

Table 1

Basic clinical characteristics

Characteristics Total (N=2,311) Dead (N=560) Alive (N=1,751) P value
Race, n (%) <0.001
   White 1,628 (70.4) 369 (65.9) 1,259 (71.9)
   Black 453 (19.6) 153 (27.3) 300 (17.1)
   Others 230 (10.0) 38 (6.8) 192 (11.0)
Primary site, n (%) 0.002
   Central portion 46 (2.0) 15 (2.7) 31 (1.8)
   Upper-inner quadrant 312 (13.5) 62 (11.1) 250 (14.3)
   Lower-inner quadrant 142 (6.1) 40 (7.1) 102 (5.8)
   Upper-outer quadrant 892 (38.6) 201 (35.9) 691 (39.5)
   Lower-outer quadrant 190 (8.2) 35 (6.3) 155 (8.9)
   Axillary tail 18 (0.8) 5 (0.9) 13 (0.7)
   Overlapping lesion 475 (20.6) 123 (22.0) 352 (20.1)
   Breast, NOS 236 (10.2) 79 (14.1) 157 (9.0)
T stage, n (%) <0.001
   T1 665 (28.8) 98 (17.5) 567 (32.4)
   T2 1,206 (52.2) 254 (45.4) 952 (54.4)
   T3 300 (13.0) 123 (22.0) 177 (10.1)
   T4 140 (6.1) 85 (15.2) 55 (3.1)
N stage, n (%) <0.001
   N0 1,303 (56.4) 169 (30.2) 1,134 (64.8)
   N1 729 (31.5) 230 (41.1) 499 (28.5)
   N2 142 (6.1) 76 (13.6) 66 (3.8)
   N3 137 (5.9) 85 (15.2) 52 (3.0)
Surgery of primary site, n (%) <0.001
   No surgery 155 (6.7) 81 (14.5) 74 (4.2)
   Breast-conserving surgery 675 (29.2) 122 (21.8) 553 (31.6)
   Total mastectomy 901 (39.0) 162 (28.9) 739 (42.2)
   Radical mastectomy 580 (25.1) 195 (34.8) 385 (22.0)
Radiation, n (%) <0.001
   No 1,212 (52.4) 257 (45.9) 955 (54.5)
   Yes 1,099 (47.6) 303 (54.1) 796 (45.5)
Chemotherapy, n (%) 0.335
   No 194 (8.4) 41 (7.3) 153 (8.7)
   Yes 2,117 (91.6) 519 (92.7) 1,598 (91.3)
Bone, n (%) <0.001
   No 2,263 (97.9) 516 (92.1) 1,747 (99.8)
   Yes 48 (2.1) 44 (7.9) 4 (0.2)
Brain, n (%) <0.001
   No 2,302 (99.6) 552 (98.6) 1,750 (99.9)
   Yes 9 (0.4) 8 (1.4) 1 (0.1)
Liver, n (%) <0.001
   No 2,278 (98.6) 530 (94.6) 1,748 (99.8)
   Yes 33 (1.4) 30 (5.4) 3 (0.2)
Lung, n (%) <0.001
   No 2,266 (98.1) 521 (93.0) 1,745 (99.7)
   Yes 45 (1.9) 39 (7.0) 6 (0.3)

N, node; NOS, not otherwise specified; T, tumor.

Radiotherapy was administered in 47.6% of patients, with a significantly higher rate in the deceased group (P<0.001). Chemotherapy was used in 91.6% of patients, with no significant difference between groups (P=0.33). The overall rates of distant metastasis were 2.1% for bone, 1.9% for lung, 1.4% for liver, and 0.4% for brain.

To verify the generalization performance of the model, patients were randomly divided into a training set (n=1,733) and a validation set (n=578) at a ratio of 3:1 using stratified randomization. The distribution of baseline characteristics between the two groups is shown in Table 2. Statistical analysis revealed no significant differences (P>0.01) between the training set and the validation set in key indicators, providing a reliable data foundation for subsequent model construction and validation.

Table 2

Baseline clinical characteristics of the training and validation sets

Characteristics Training (N=1,733) Testing (N=578) P value
Race, n (%) 0.74
   White 1,222 (70.5) 406 (70.2)
   Black 343 (19.8) 110 (19.0)
   Others 168 (9.7) 62 (10.7)
Primary site, n (%) 0.35
   Central portion 39 (2.3) 7 (1.2)
   Upper-inner quadrant 232 (13.4) 80 (13.8)
   Lower-inner quadrant 111 (6.4) 31 (5.4)
   Upper-outer quadrant 667 (38.5) 225 (38.9)
   Lower-outer quadrant 143 (8.3) 47 (8.1)
   Axillary tail 13 (0.8) 5 (0.9)
   Overlapping lesion 364 (21.0) 111 (19.2)
   Breast, NOS 164 (9.5) 72 (12.5)
T stage, n (%) 0.15
   T1 490 (28.3) 175 (30.3)
   T2 927 (53.5) 279 (48.3)
   T3 214 (12.3) 86 (14.9)
   T4 102 (5.9) 38 (6.6)
N stage, n (%) 0.90
   N0 981 (56.6) 322 (55.7)
   N1 543 (31.3) 186 (32.2)
   N2 104 (6.0) 38 (6.6)
   N3 105 (6.1) 32 (5.5)
Surgery of primary site, n (%) 0.73
   No surgery 112 (6.5) 43 (7.4)
   Breast-conserving surgery 502 (29.0) 173 (29.9)
   Total mastectomy 685 (39.5) 216 (37.4)
   Radical mastectomy 434 (25.0) 146 (25.3)
Radiation, n (%) 0.66
   No 914 (52.7) 298 (51.6)
   Yes 819 (47.3) 280 (48.4)
Chemotherapy, n (%) 0.49
   No 141 (8.1) 53 (9.2)
   Yes 1,592 (91.9) 525 (90.8)
Bone, n (%) >0.99
   No 1,697 (97.9) 566 (97.9)
   Yes 36 (2.1) 12 (2.1)
Brain, n (%) 0.56
   No 1,725 (99.5) 577 (99.8)
   Yes 8 (0.5) 1 (0.2)
Liver, n (%) 0.13
   No 1,704 (98.3) 574 (99.3)
   Yes 29 (1.7) 4 (0.7)
Lung, n (%) 0.43
   No 1,702 (98.2) 564 (97.6)
   Yes 31 (1.8) 14 (2.4)

N, node; NOS, not otherwise specified; T, tumor.

Variable screening

To identify the key features influencing CSS in young women with triple-negative invasive ductal breast cancer, we adopted Boruta analysis to evaluate feature importance, with CSS set as the target variable. The results are presented in Figure 1 (Boruta feature importance).

Figure 1 Boruta feature importance rankings (with cancer-specific survival as the target variable). The error bars in the figure represent the variation in feature importance across repeated evaluations. The N stage, T stage, and distant metastasis-related features consistently showed high scores, indicating stable predictive value for cancer-specific survival. N, node; T, tumor.

In this analysis, the “shadow” features (e.g., shadowWin, shadowMean) were used as the reference baseline to judge the significance of the real features. A real feature with an importance score exceeding the baseline of the shadow features was considered to have a significant predictive effect. As shown in the figure, the features of N stage (N), T stage (T), distant metastasis (including bone, lung, and liver metastasis), and primary surgical approach (Surg_Prim) exhibited markedly higher importance scores (represented by green and dark-colored bars), indicating that they were significant predictors of CSS. In contrast, the importance scores of features such as primary tumor site (most subclasses) and race were relatively low, being close to or even lower than the baseline of the shadow features, suggesting weak predictive correlations with CSS.

Predictive performance and visualization analysis of ML models

In this study, eight ML models were constructed, and a multi-dimensional evaluation of their predictive performance for CSS in young women with TN-IDC was conducted. The core results are summarized in Table 3 and Figure 2.

Table 3

Predictive performance of machine learning models with 95% CI

Model Accuracy Kap Sensitivity Specificity PPV NPV MCC J_index Bal_accuracy
Logistic 0.689 (0.66–0.717) 0.298 (0.27–0.327) 0.696 (0.668–0.725) 0.664 (0.635–0.694) 0.866 (0.845–0.888) 0.412 (0.381–0.442) 0.317 (0.288–0.345) 0.361 (0.331–0.39) 0.68 (0.651–0.709)
enet 0.749 (0.722–0.776) 0.35 (0.32–0.379) 0.811 (0.786–0.835) 0.557 (0.526–0.588) 0.851 (0.829–0.873) 0.484 (0.453–0.515) 0.351 (0.322–0.381) 0.368 (0.338–0.398) 0.684 (0.655–0.713)
dt 0.765 (0.738–0.791) 0.241 (0.214–0.268) 0.918 (0.901–0.935) 0.286 (0.258–0.314) 0.801 (0.776–0.826) 0.526 (0.495–0.557) 0.258 (0.231–0.285) 0.204 (0.179–0.228) 0.602 (0.571–0.632)
rf 0.756 (0.729–0.783) 0.356 (0.326–0.385) 0.824 (0.801–0.848) 0.543 (0.512–0.574) 0.849 (0.827–0.872) 0.497 (0.466–0.528) 0.356 (0.327–0.386) 0.367 (0.337–0.397) 0.684 (0.655–0.712)
xgboost 0.701 (0.672–0.729) 0.319 (0.29–0.348) 0.71 (0.682–0.738) 0.671 (0.642–0.701) 0.871 (0.85–0.892) 0.425 (0.395–0.456) 0.336 (0.307–0.366) 0.381 (0.351–0.412) 0.691 (0.662–0.719)
rsvm 0.772 (0.746–0.798) 0.356 (0.326–0.386) 0.865 (0.844–0.886) 0.479 (0.448–0.51) 0.838 (0.816–0.861) 0.532 (0.501–0.563) 0.357 (0.327–0.387) 0.344 (0.314–0.373) 0.672 (0.643–0.701)
knn 0.77 (0.744–0.796) 0.362 (0.333–0.392) 0.856 (0.834–0.878) 0.5 (0.469–0.531) 0.843 (0.82–0.865) 0.526 (0.495–0.557) 0.363 (0.333–0.392) 0.356 (0.326–0.386) 0.678 (0.649–0.707)
MLP 0.711 (0.683–0.739) 0.37 (0.341–0.4) 0.771 (0.745–0.797) 0.692 (0.663–0.72) 0.444 (0.414–0.475) 0.904 (0.886–0.923) 0.402 (0.372–0.432) 0.463 (0.432–0.494) 0.732 (0.704–0.759)

Accuracy refers to the overall prediction accuracy, reflecting the proportion of cases correctly predicted by the model; Kap denotes Cohen’s Kappa coefficient, which measures the consistency between the model’s predicted results and the actual clinical outcomes; Sensitivity reflects the model’s ability to identify patients with poor prognosis; Specificity indicates the model’s ability to identify patients with favorable prognosis; PPV is the positive predictive value, representing the credibility of the model when it indicates a poor prognosis; NPV is the negative predictive value, reflecting the credibility of the model when it indicates a favorable prognosis; MCC refers to the Matthews correlation coefficient, a balanced index that comprehensively integrates all types of prediction results; J_index denotes Youden’s index, which reflects the balance between the model’s sensitivity and specificity; Bal_accuracy is the balanced accuracy, which is suitable for imbalanced classification data and reflects the model’s balanced predictive ability for the two types of outcomes. CI, confidence interval; dt, decision tree; enet, elastic net; knn, k-nearest neighbors; logistic, logistic regression; MLP, multilayer perceptron; rf, random forest; rsvm, radial support vector machine; xgboost, extreme gradient boosting.

Figure 2 Multi-metric performance comparison of machine learning models across evaluation indices. Accuracy refers to the overall prediction accuracy, reflecting the proportion of cases correctly predicted by the model; Kap denotes Cohen’s Kappa coefficient, which measures the consistency between the model’s predicted results and the actual clinical outcomes; Sensitivity reflects the model’s ability to identify patients with poor prognosis; Specificity indicates the model’s ability to identify patients with favorable prognosis; PPV is the positive predictive value, representing the credibility of the model when it indicates a poor prognosis; NPV is the negative predictive value, reflecting the credibility of the model when it indicates a favorable prognosis; MCC refers to the Matthews correlation coefficient, a balanced index that comprehensively integrates all types of prediction results; J_index denotes Youden’s index, which reflects the balance between the model’s sensitivity and specificity; Bal_accuracy is the balanced accuracy, which is suitable for imbalanced classification data and reflects the model’s balanced predictive ability for the two types of outcomes. CI, confidence interval; dt, decision tree; enet, elastic net; knn, k-nearest neighbors; logistic, logistic regression; MLP, multilayer perceptron; rf, random forest; rsvm, radial support vector machine; xgboost, extreme gradient boosting.

Comparison of core model performance: The MLP model exhibited the optimal overall performance, achieving the highest values among all models for sensitivity [0.771, 95% confidence interval (CI): 0.683–0.739], MCC (0.402, 95% CI: 0.372–0.423), and balanced accuracy (0.732, 95% CI: 0.704–0.759). These results indicate that the MLP model has significant advantages in identifying high-risk patients, quantifying risk probabilities, and generating balanced predictions for the two outcome categories. The rsvm showed outstanding performance in accuracy (0.772), PPV =0.838, and Kap =0.356, demonstrating its superiority in improving the credibility of predictions. The dt model achieved the highest sensitivity (0.918) but had a relatively low specificity (0.286), making it more suitable for scenarios with low misdiagnosis tolerance. The logistic regression model yielded moderate performance across most metrics, with relatively high PPV (0.866) but low sensitivity (0.696) and kappa (0.298). The xgboost model delivered balanced performance, with a competitive balanced accuracy of 0.691, though it did not rank highest in any single metric.

ROC curve analysis of the eight predictive models (Figure 3) revealed distinct differences in their performance for predicting CSS among young women with TN-IDC. As illustrated in Figure 3, the MLP model exhibited the most robust discriminative ability (AUC =0.782, 95% CI: 0.737–0.828), with its ROC curve positioned farthest toward the upper-left corner. Notably, logistic regression (0.736, 95% CI: 0.687–0.784), enet (0.738, 95% CI: 0.688–0.788), and xgboost (0.741, 95% CI: 0.694–0.789) also demonstrated favorable discriminative performance, whereas the knn (0.698, 95% CI: 0.649–0.747) and dt (0.393, 95% CI: 0.353–0.434) models showed significantly insufficient efficiency. The dt model generates only a small number of discrete probability values, determined by the proportion of positive samples in leaf nodes. This produces a stepwise ROC curve rather than a smooth, continuous one.

Figure 3 Receiver operating characteristic curves of various machine learning models. AUC, area under the curve; CI, confidence interval; dt, decision tree; enet, elastic net; knn, k-nearest neighbors; logistic, logistic regression; MLP, multilayer perceptron; rf, random forest; rsvm, radial support vector machine; xgboost, extreme gradient boosting.

DeLong’s test results indicated that the AUC of the MLP model was significantly higher than that of knn (P=0.01) and rsvm (P=0.02), with statistically significant differences. In contrast, no significant AUC differences were observed between MLP and enet (P=0.19), logistic regression (P=0.17), rf (P=0.10), or xgboost (P=0.22) (all P>0.05). Nevertheless, MLP remained numerically superior across all models, further validating its reliability and comparative advantage in predicting CSS for this specific patient cohort (Figure S1).

Calibration curves

To evaluate the reliability of each model in predicting the prognosis of young women with triple-negative invasive ductal breast cancer, calibration curves of all models on the test set were plotted (as shown in Figure 4; the dashed line represents the ideal calibration line where the predicted probability is equal to the actual event rate). Among all the models, the MLP model exhibited the highest degree of fit between its curve and the ideal calibration line: its red curve was close to the dashed line across the entire range of predicted probabilities, indicating that the risk probabilities output by this model were highly consistent with the actual outcomes. The xgboost model also showed good performance, with its curve gradually approaching the ideal line as the predicted probability increased. In contrast, the curve of the knn model deviated significantly from the ideal line, especially in the middle range of predicted probabilities, suggesting a large discrepancy between its predicted probabilities and the actual event rates. Although the other models (e.g., dt, enet, logistic, rf, rsvm) showed a certain degree of fit with the ideal line in some intervals, they exhibited obvious fluctuations in specific probability segments, resulting in lower reliability than the MLP and xgboost models.

Figure 4 Calibration curves indicating the alignment of predicted probabilities with observed outcomes. dt, decision tree; enet, elastic net; knn, k-nearest neighbors; logistic, logistic regression; MLP, multilayer perceptron; rf, random forest; rsvm, radial support vector machine; xgboost, extreme gradient boosting.

MLP artificial neural network modeling

Figure 5 presents the architecture of the artificial neural network (MLP) constructed in this study. Adopting a feedforward design, the network consists of three layers: an input layer, a hidden layer, and an output layer. The input layer (labeled with features including Race_X2, T_X3, N_X1, and Bone_X1) comprises 16 nodes, with each node corresponding to a preprocessed clinical feature (e.g., race subclass, T/N stage subclass, distant metastasis status, and surgical approach). The hidden layer (marked as H1 to H24) contains 24 nodes, which are responsible for performing a nonlinear transformation on the input features. The output layer (O1) includes 1 node, corresponding to the target variable (prognostic outcome, labeled as “.y”). In this architecture, the connections between the input layer and the hidden layer (Region B1) represent the weight assignment from the input features to the hidden nodes, while the connections between the hidden layer and the output layer (Region B2) represent the weight assignment from the hidden nodes to the output. This multi-layer mapping enables the model to capture the complex nonlinear relationships between clinical features and prognostic outcomes in young women with TN-IDC.

Figure 5 Architecture of the MLP neural network model for prognostic prediction in young women with TN-IDC. MLP, multilayer perceptron; N, node; T, tumor; TN-IDC, triple-negative invasive ductal breast cancer.

Feature importance in MLP models

To clarify the contribution of each clinical feature to the prognosis prediction of young women with triple-negative invasive ductal breast cancer, we analyzed the feature importance (Figure 6), SHAP values (Figure 7) of the model and SHAP waterfall plot for a single-sample (Figure S2).

Figure 6 Feature importance ranking of the predictive model (sorted by importance score). Bar chart is sorted by the magnitude of feature importance (bar length). N_X3 (a subclass of N stage) has the largest importance value, followed by Bone_X1 (bone metastasis) and N_X2, indicating that these are the most critical predictors of prognosis. In contrast, features such as Surg.Prim.Site_X1 have shorter bars, representing relatively weak predictive effects. N, node; T, tumor.
Figure 7 SHAP summary plot for the MLP prognostic prediction model. The plot illustrates the contribution of each input feature to the model’s predictions of adverse prognosis in young TN-IDC patients. The horizontal axis represents the SHAP value: a positive SHAP value corresponds to an increased risk of adverse prognosis, while a negative SHAP value corresponds to a decreased risk of adverse prognosis. The vertical axis ranks feature by their overall importance to the model (from highest to lowest). The color of each dot indicates the feature value (blue = low, red = high), and the distribution of dots reflects the range of SHAP values for each feature across the entire cohort. MLP, multilayer perceptron; N, node; SHAP, Shapley Additive Explanations; TN-IDC, triple-negative invasive ductal breast cancer.

Figure 7 shows the degree of influence of each feature on the prediction results: the N stage subclass N_X3 has the highest importance, followed by bone metastasis (Bone_X1) and the N stage subclass N_X2. In contrast, features such as the surgical primary site subclass Surg.Prim.Site_X1 have relatively low importance, indicating that their roles in the model’s decision-making process are weak.

As shown in the SHAP summary plot (Figure 7), we quantified the contribution of each feature to the MLP model’s prediction of CSS. The vertical axis ranks features by importance, while the horizontal axis represents the SHAP value, with the color gradient denoting feature values from low (blue) to high (red). A positive SHAP value indicates an increased risk of adverse CSS, whereas a negative value indicates a reduced risk. Consistent with clinical expectations, high-risk factors such as advanced nodal stage (N_X3) and bone metastasis (Bone_X1) predominantly showed positive SHAP values at higher feature levels (red points), confirming their role in elevating the risk of cancer-specific death. Conversely, protective factors including certain surgical approaches (e.g., Surg.Prim.Site_X2) were associated with negative SHAP values at higher levels, indicating a reduced risk of adverse outcomes. Other variables (e.g., Race_X2, T_X2) showed narrower distributions and weaker effects on model predictions. This analysis enhances the interpretability of the MLP model and aligns with known prognostic patterns in young women with TN-IDC.


Discussion

In this study, we focused on TN-IDC in young women—a subtype characterized by high aggressiveness, and established eight ML prediction models with comprehensive validation through the integration of multi-faceted clinical data. By identifying the key determinants of prognosis, our findings provide an innovative and quantitative framework for predicting CSS and individualized treatment in clinical practice.

Through Boruta feature analysis, this study identified the N stage, T stage, distant metastasis (bone, lung, and liver), and surgical approach as the core prognostic predictors for TN-IDC in young women, with the N stage subclass (N_X3) and bone metastasis (Bone_X1) exerting the most prominent effects. This finding is highly consistent with the biological characteristics of TN-IDC: lymph node metastasis directly reflects tumor invasiveness, where a higher N stage indicates stronger metastatic potential of tumor cells. Moreover, the breast tissue of young patients exhibits high proliferative activity, rendering them more susceptible to lymphatic distant metastasis, which aligns with the conclusions of previous studies that “patients with N2/N3 stage disease have a significantly increased risk of recurrence” (19). As one of the most common distant metastatic sites in TN-IDC, bone metastasis exhibited a significantly higher incidence in the deceased group than in the survivor group. Its occurrence severely compromises patients’ quality of life and overall survival, which explains its critical role as a core prognostic predictor (20). Surgical approach also demonstrated distinct clinical value for prognosis. Specifically, the deceased group had a markedly higher proportion of patients who received no surgery or radical mastectomy, whereas the survivor group showed a more favorable distribution of breast-conserving surgery and total mastectomy, with statistically significant differences between the two groups. This observation indicates that standardized surgical treatment is a key determinant of improved prognosis. Notably, patients who did not undergo surgery faced a higher risk of death, primarily due to rapid disease progression, poor baseline health status, or refusal of treatment. Although radical mastectomy enables complete resection of the lesion, it is associated with greater surgical trauma, which may impair postoperative recovery and subsequent treatment tolerance (21). The high proportion of radical mastectomy and non-surgical cases in the deceased group indirectly reflects the more advanced disease status of these patients. This finding not only provides direct evidence for clinical treatment decision-making but also is consistent with the results of variable screening, which confirmed the stable predictive efficacy of surgical approach as a prognostic feature.

Following the identification of core prognostic predictors via Boruta-SHAP feature selection, this study further compared the performance of eight ML models. The MLP model exhibited the optimal comprehensive predictive performance, with leading values across all key metrics. Additionally, its calibration curve showed the highest degree of fit to the ideal line. These results indicate that the MLP model not only possesses robust prognostic discrimination capability but also enables accurate quantification of risk probabilities, which are highly consistent with actual clinical outcomes. This superiority stems from the multi-layer neural network architecture of the MLP model, which can effectively capture the complex non-linear relationships between clinical features and prognosis, making it particularly suitable for highly heterogeneous tumor subtypes such as TN-IDC (22-24). The rsvm model demonstrated outstanding performance in terms of accuracy, PPV, and Cohen’s kappa coefficient. It offers advantages in enhancing the credibility of predictions and is thus well-suited for clinical decision-making scenarios with high requirements for result reliability, such as the optimization of treatment regimens for high-risk patients. In contrast, the dt model showcased its strengths with excellent specificity. Despite its relatively low sensitivity (0.298), its strong interpretability allows clinicians to quickly understand the underlying prediction logic and apply it without the need for advanced algorithmic expertise. This makes it ideal for primary medical institutions or rapid risk screening settings, where it can effectively reduce the misdiagnosis rate of low-risk patients (25,26). By comparison, the knn model performed the poorest, with an accuracy of only 0.514 and an AUC of 0.54, indicating a discrimination ability close to random chance. This suboptimal performance may be attributed to the high dimensionality, class imbalance, and complex inter-feature correlations inherent in TN-IDC data, which make it difficult for simple distance-based measurement methods to capture key predictive patterns. These findings suggest that the knn model is not suitable for prognostic prediction in this disease setting. In terms of clinical benchmarking, the MLP model achieved an AUC of 0.782, representing moderate-to-good discriminative ability for CSS in young women with TN-IDC. This performance is comparable to or even superior to that of conventional clinical prognostic systems, including the AJCC staging system, which primarily rely on tumor stage and basic pathological features. Traditional tools often fail to integrate comprehensive information, including metastatic status and surgical modalities, simultaneously. In contrast, our ML model incorporates multi-dimensional clinical variables and captures non-linear relationships, thereby providing more refined risk stratification. Although the model does not completely replace established clinical tools, it offers a supplementary and interpretable strategy to improve individualized prognostic evaluation for this high-risk population. Furthermore, although the xgboost, logistic regression, elastic net, and random forest models did not achieve the highest performance, they all exhibited a certain level of predictive efficacy (AUC range: 0.67–0.75). These models can be used as alternatives for ensemble prediction, thereby further improving the reliability of the results.

The core value of the MLP model elucidated by the SHAP framework is its ability to overcome the limitations of traditional prognostic evaluation and provide an accurate and interpretable tool for predicting CSS in young women with TN-IDC. As a widely recognized model-agnostic interpretable artificial intelligence method, SHAP quantifies the direction and magnitude of each feature’s contribution to model prediction, thereby improving the transparency and clinical applicability of artificial intelligence-based predictive models (27). Feature importance ranking demonstrated that N stage subgroup (N_X3) was the most influential prognostic factor, followed by bone metastasis (Bone_X1) and N_X2, while the contribution of surgical approach subgroups (e.g., Surg.Prim.Site_X1) was relatively low. This ranking was highly consistent with the results of Boruta feature selection.

Furthermore, the SHAP summary plot clearly illustrated the directional effect of each feature on CSS. A positive SHAP value indicates an increased risk of cancer-specific death, whereas a negative SHAP value indicates a reduced risk. High feature values (red) for N stage-related variables (e.g., N_X1, N_X3) were concentrated in the positive SHAP region, confirming that advanced lymph node metastasis significantly elevates the risk of poor CSS. By contrast, high feature values (red) for Surg.Prim.Site_X2 were mainly distributed in the negative SHAP region, indicating that appropriate surgical (e.g., breast-conserving surgery or total mastectomy) treatment may reduce the risk of adverse CSS outcomes. This finding corroborates the clinical consensus that standardized surgical treatment improves the prognosis of TN-IDC patients (28-30). In addition, lung metastasis (Lung_X1) was also identified as an important risk factor, and its effect magnitude was consistent with the feature importance ranking. This observation underscores the prognostic heterogeneity of distant metastatic sites in TN-IDC and highlights the importance of identifying distant metastasis for risk stratification.

The SHAP waterfall plot for the individual sample further illustrates the personalized prognostic contribution of each clinical and pathological feature, thereby providing transparent and intuitive support for individualized risk assessment in young women with TN-IDC. As shown in the waterfall plot (sample 99, Figure S2), the final predicted probability of CSS was jointly determined by multiple key variables, including tumor stage (T_X4), lymph node stage (N_X1), surgical approach (Surg.Prim.Site_X3), and race (Race_X2). Among these, advanced tumor stage (T_X4) and lymph node metastasis (N_X1) represented the major high-risk features that elevated the prognostic risk, while certain surgical approaches and demographic variables exerted mild protective effects. This personalized interpretation demonstrates that the MLP model can quantify the unique contribution of each clinical feature in individual patients, which helps clinicians identify the most influential determinants of prognosis for patients with TN-IDC. Such feature-specific risk decomposition allows for refined risk communication, more targeted follow-up arrangements, and better alignment with real-world clinical decision-making. For patients with prominent high-risk features such as advanced T stage and lymph node metastasis, closer monitoring and more intensive comprehensive therapy may be warranted. For those dominated by protective factors, appropriately reduced intensity of follow-up may be considered to avoid redundant medical interventions. In summary, single-sample SHAP analysis enhances the clinical interpretability of the model by revealing the individualized mechanism of prognostic prediction, thereby facilitating personalized risk stratification and clinical management for young women with TN-IDC.

This study has certain limitations. First, although class weight adjustment was used to alleviate training bias, external validation with more balanced populations is warranted. Second, the exclusion of molecular markers such as BRCA1/2 gene mutations may compromise the predictive efficacy of the model. Third, the failure to perform a stratified analysis of the efficacy of different treatment regimens hinders the direct guidance for precise treatment selection. Future research can further optimize model performance and promote the development of personalized precision medicine by conducting multi-center prospective studies, supplementing molecular and lifestyle indicators, and constructing full-process predictive models integrated with treatment response data.


Conclusions

This study identified the core prognostic features of TN-IDC in young women through the Boruta-SHAP interpretable ML framework. The MLP model functions as a supplementary tool to support clinically precise risk stratification, individualized treatment decision-making, and structured follow-up management. Clinical implementation of these prognostic features, in conjunction with the model, may help optimize resource distribution, reduce inappropriate treatment, and ultimately improve survival outcomes for this population.


Acknowledgments

None.


Footnote

Reporting Checklist: The authors have completed the TRIPOD reporting checklist. Available at https://gs.amegroups.com/article/view/10.21037/gs-2026-1-0091/rc

Peer Review File: Available at https://gs.amegroups.com/article/view/10.21037/gs-2026-1-0091/prf

Funding: None.

Conflicts of Interest: All authors have completed the ICMJE uniform disclosure form (available at https://gs.amegroups.com/article/view/10.21037/gs-2026-1-0091/coif). The authors have no conflicts of interest to declare.

Ethical Statement: The authors are accountable for all aspects of the work in ensuring that questions related to the accuracy or integrity of any part of the work are appropriately investigated and resolved. The study was conducted in accordance with the Declaration of Helsinki and its subsequent amendments.

Open Access Statement: This is an Open Access article distributed in accordance with the Creative Commons Attribution-NonCommercial-NoDerivs 4.0 International License (CC BY-NC-ND 4.0), which permits the non-commercial replication and distribution of the article with the strict proviso that no changes or edits are made and the original work is properly cited (including links to both the formal publication through the relevant DOI and the license). See: https://creativecommons.org/licenses/by-nc-nd/4.0/.


References

  1. The global, regional, and national burden of cancer, 1990-2023, with forecasts to 2050: a systematic analysis for the Global Burden of Disease Study 2023. Lancet 2025;406:1565-86.
  2. Raman R, Debata S, Govindarajan T, et al. Targeting Triple-Negative Breast Cancer: Resistance Mechanisms and Therapeutic Advancements. Cancer Med 2025;14:e70803. [Crossref] [PubMed]
  3. Early Breast Cancer Trialists’ Collaborative Group. Reductions in recurrence in women with early breast cancer entering clinical trials between 1990 and 2009: a pooled analysis of 155 746 women in 151 trials. Lancet 2024;404:1407-18. Erratum in: Lancet 2025;405:1816.
  4. Jie H, Ma W, Huang C. Diagnosis, Prognosis, and Treatment of Triple-Negative Breast Cancer: A Review. Breast Cancer (Dove Med Press) 2025;17:265-74. [Crossref] [PubMed]
  5. Yoon J, Knapp G, Quan ML, et al. Cancer-Specific Outcomes in the Elderly with Triple-Negative Breast Cancer: A Systematic Review. Curr Oncol 2021;28:2337-45. [Crossref] [PubMed]
  6. Ahmed B, Al-Khames Aga Q, Cheung KL, et al. Treatment strategies for triple-negative primary breast cancer in older women: a systematic review. JNCI Cancer Spectr 2025;9:pkaf049. [Crossref] [PubMed]
  7. Pohl-Rescigno E, Hauke J, Loibl S, et al. Association of Germline Variant Status With Therapy Response in High-risk Early-Stage Breast Cancer: A Secondary Analysis of the GeparOcto Randomized Clinical Trial. JAMA Oncol 2020;6:744-8. [Crossref] [PubMed]
  8. Gheller A, Tuttle R, Oxenberg J. Breast Cancer Risk, Screening, and Risk Reduction in Young Females. Am Surg 2025;91:1178-87. [Crossref] [PubMed]
  9. Bai X, Ni J, Beretov J, et al. Triple-negative breast cancer therapeutic resistance: Where is the Achilles’ heel? Cancer Lett 2021;497:100-11. [Crossref] [PubMed]
  10. Aguiar-Ibáñez R, Mbous YPV, Sharma S, et al. Assessing the clinical, humanistic, and economic impact of early cancer diagnosis: a systematic literature review. Front Oncol 2025;15:1546447. [Crossref] [PubMed]
  11. Zhang Y, Ji Y, Liu S, et al. Global burden of female breast cancer: new estimates in 2022, temporal trend and future projections up to 2050 based on the latest release from GLOBOCAN. J Natl Cancer Cent 2025;5:287-96. [Crossref] [PubMed]
  12. Yang T, Zhu Z, Shi J, et al. Association among financial toxicity, depression and fear of cancer recurrence in young breast cancer patient-family caregiver dyads: an actor-partner interdependence mediation model. BMC Psychiatry 2025;25:97. [Crossref] [PubMed]
  13. Rockall AG, Chiu SM, Aboagye EO, et al. Artificial intelligence in women’s cancers: innovation and challenges in clinical translation. Lancet Digit Health 2025;7:100940. [Crossref] [PubMed]
  14. Nguyen THT, Jeon S, Yoon J, et al. Global mapping of artificial intelligence applications in breast cancer from 1988-2024: a machine learning approach. Breast Cancer 2025;32:1244-54. [Crossref] [PubMed]
  15. Peng H, Wang Y, Tang X, et al. Adjustable spot wide-field Raman spectroscopy combined with machine learning for accurate classification of breast cancer cells. Talanta 2026;298:128862. [Crossref] [PubMed]
  16. Valizadeh P, Jannatdoust P, Moradi N, et al. Ultrasound-based machine learning models for predicting response to neoadjuvant chemotherapy in breast cancer: A meta-analysis. Clin Imaging 2025;125:110574. [Crossref] [PubMed]
  17. Bunte K, Ituarte B, Warikoo G, et al. Regional disparities in incidence and outcomes of invasive ductal carcinoma of the breast among Asian and Pacific Islander women in the United States. Cancer Epidemiol 2025;97:102861. [Crossref] [PubMed]
  18. Ejiyi CJ, Qin Z, Ukwuoma CC, et al. Comparative performance analysis of Boruta, SHAP, and Borutashap for disease diagnosis: A study with multiple machine learning algorithms. Network 2025;36:507-44. [Crossref] [PubMed]
  19. Ji K, Zhao Z, Sameni M, et al. Modeling Tumor: Lymphatic Interactions in Lymphatic Metastasis of Triple Negative Breast Cancer. Cancers (Basel) 2021;13:6044. [Crossref] [PubMed]
  20. Zari DS, Novak R, Haviv O, et al. Anatomical distribution of bone metastases in stage IV breast cancer: According to histological subtype. World J Clin Oncol 2025;16:110087. [Crossref] [PubMed]
  21. Sekiguchi K, Kawamori J, Yamauchi H. Breast reconstruction and postmastectomy radiotherapy: complications by type and timing and other problems in radiation oncology. Breast Cancer 2017;24:511-20. [Crossref] [PubMed]
  22. Gao Y, Liu L, Wang S, et al. SEER-based machine learning prediction of bone metastasis in breast cancer: model development and validation. Gland Surg 2025;14:1366-78. [Crossref] [PubMed]
  23. Srinivasu PN, Jaya Lakshmi G, Gudipalli A, et al. XAI-driven CatBoost multi-layer perceptron neural network for analyzing breast cancer. Sci Rep 2024;14:28674. [Crossref] [PubMed]
  24. Lee EG, Han J, Lee S, et al. Effectiveness of Multi-Layer Perceptron-Based Binary Classification Neural Network in Detecting Breast Cancer Through Nine Human Serum Protein Markers. Cancers (Basel) 2025;17:2832. [Crossref] [PubMed]
  25. Wang X, Zhang D. Choices of medical institutions and associated factors in older patients with multimorbidity in stabilization period in China: A study based on logistic regression and decision tree model. Health Care Sci 2023;2:359-69. [Crossref] [PubMed]
  26. Zhang L, Liu H, Zou Z, et al. Shared-care models are highly effective and cost-effective for managing chronic hepatitis B in China: reinterpreting the primary care and specialty divide. Lancet Reg Health West Pac 2023;35:100737. [Crossref] [PubMed]
  27. Ghasemi A, Hashtarkhani S, Schwartz DL, et al. Explainable artificial intelligence in breast cancer detection and risk prediction: A systematic scoping review. Cancer Innov 2024;3:e136. [Crossref] [PubMed]
  28. Li Y, Zhang H, Merkher Y, et al. Recent advances in therapeutic strategies for triple-negative breast cancer. J Hematol Oncol 2022;15:121. [Crossref] [PubMed]
  29. Park KU, Somerfield MR, Anne N, et al. Sentinel Lymph Node Biopsy in Early-Stage Breast Cancer: ASCO Guideline Update. J Clin Oncol 2025;43:1720-41. [Crossref] [PubMed]
  30. Magnoni F, Alessandrini S, Alberti L, et al. Breast Cancer Surgery: New Issues. Curr Oncol 2021;28:4053-66. [Crossref] [PubMed]
Cite this article as: Luo L, Li S, Ji Q. Interpretable machine learning for prognostic prediction in young women with triple-negative invasive ductal breast cancer: a Boruta-SHAP integrated approach. Gland Surg 2026;15(5):120. doi: 10.21037/gs-2026-1-0091

Download Citation