Development and validation of a machine learning model for predicting liver metastasis in pancreatic cancer: a retrospective cohort study
Original Article

Development and validation of a machine learning model for predicting liver metastasis in pancreatic cancer: a retrospective cohort study

Ziang Chen1,2#, Peng Song1# ORCID logo, Hong Fang1#, Wencong Tian1, Jia Zhao1, Yanhong Liu1, Chuntao Wang1, Yongjie Zhao1, Lei Cao1 ORCID logo

1Department of General Surgery, Tianjin Union Medical Center, The First Affiliated Hospital of Nankai University, Tianjin, China; 2School of Medicine, Nankai University, Tianjin, China

Contributions: (I) Conception and design: Z Chen, P Song; (II) Administrative support: Y Zhao, L Cao; (III) Provision of study materials or patients: H Fang, W Tian; (IV) Collection and assembly of data: P Song, J Zhao; (V) Data analysis and interpretation: Y Liu, C Wang; (VI) Manuscript writing: All authors; (VII) Final approval of manuscript: All authors.

#These authors contributed equally to this work.

Correspondence to: Yongjie Zhao, MM; Lei Cao, MD. Department of General Surgery, Tianjin Union Medical Center, The First Affiliated Hospital of Nankai University, No. 190, Jieyuan Road, Hongqiao District, Tianjin 300121, China. Email: zhaoyongjieky@126.com; caolei@umc.net.cn.

Background: Liver metastasis (LM) is the leading cause of treatment failure and poor survival rates in pancreatic cancer (PC). Early identification of patients at high risk for LM is crucial for optimizing clinical management. This study aims to develop and validate a machine learning (ML)-based predictive model to accurately identify the risk of LM in PC patients and to determine the core risk factors contributing to LM.

Methods: We collected data from eligible PC patients from the Surveillance, Epidemiology, and End Results (SEER) database (2010–2015), which was divided into a training set (70%) and a validation set (30%). Feature selection was performed using the Least Absolute Shrinkage and Selection Operator (LASSO), recursive feature elimination (RFE), and Boruta-Shap algorithms, with model parameters optimized through cross-validation. Six machine learning (ML) algorithms were developed to predict LM risk in PC patients. Model performance was evaluated using receiver operating characteristic (ROC) curve analysis, including the area under the curve (AUC), sensitivity, specificity, and negative predictive value (NPV). The best-performing model was further interpreted using SHapley Additive exPlanations (SHAP). Additionally, based on the optimal ML model, an online calculator was developed to provide personalized LM risk assessments for PC patients.

Results: A total of 12,098 eligible PC patients were registered in the SEER database, with 8,469 assigned to the training set and 3,629 to the validation set. Using a combination of LASSO, RFE, and Boruta-Shap methods, key predictors of LM in PC were identified, including surgery, radiotherapy, tumor size, age, histological grade, primary site, pathological type, T stage, lung metastasis, and bone metastasis. All six ML models performed well in the training set, with the final model selection based on the validation set results. In the validation set, the Gradient Boosting Machine (GBM) model achieved a competitive AUC value of 0.875 [95% confidence interval (CI): 0.862–0.889], with balanced sensitivity (0.851), specificity (0.776), and NPV (0.959). Furthermore, precision-recall (PR) curves, calibration curves, and decision curve analysis (DCA) consistently confirmed the superiority of the GBM model. Additionally, the SHAP framework indicated that surgery, radiotherapy, and tumor size were the primary factors influencing the predictions made by the ML model.

Conclusions: The GBM model was proven to be the most effective in predicting LM in PC patients. Its robust predictive performance and interpretability provide a reliable tool for supporting personalized clinical decision-making.

Keywords: Pancreatic cancer (PC); liver metastasis (LM); machine learning (ML); Gradient Boosting Machine (GBM); web-based calculator


Submitted Feb 04, 2026. Accepted for publication Mar 20, 2026. Published online Apr 17, 2026.

doi: 10.21037/gs-2026-1-0088


Highlight box

Key findings

• A Gradient Boosting Machine (GBM) model showed good performance for predicting liver metastasis (LM) in pancreatic cancer (PC), with strong discrimination, calibration, and clinical utility in the validation set.

• Surgery, radiotherapy, and tumor size were identified as major contributors to model prediction by SHapley Additive exPlanations (SHAP) analysis.

What is known and what is new?

• LM is common in PC and is associated with poor prognosis, but practical individualized prediction tools remain limited.

• This study integrated Least Absolute Shrinkage and Selection Operator (LASSO), Recursive Feature Elimination (RFE), and Boruta-Shap for feature selection, compared six machine learning algorithms, identified GBM as the best-performing model, and developed a web-based calculator for individualized risk assessment.

What is the implication, and what should change now?

• This model may help identify high-risk patients earlier, improve staging vigilance, and support personalized clinical decision-making.

• Further external and prospective validation is warranted before routine clinical implementation.


Introduction

Pancreatic cancer (PC) is considered one of the deadliest malignant tumors, ranking sixth in global cancer mortality and accounting for 4.8% of all cancer-related deaths (1). Despite significant advances in diagnostic and therapeutic approaches in recent years, the overall survival outcomes remain poor, with a five-year survival rate of only 13% following diagnosis (2). PC is characterized by aggressive biological features, rapid progression, and a strong tendency for local invasion and distant metastasis. Clinically, it is often diagnosed at an advanced stage, significantly limiting opportunities for curative treatment and resulting in a poor prognosis (3,4). For patients with distant metastasis, the five-year survival rate is only approximately 2.6%, and about 52% of patients are already in the distant metastatic stage at the time of diagnosis (5).

The liver is one of the most common sites of distant metastasis in pancreatic malignancies, with liver metastasis (LM) accounting for the highest proportion among metastatic patients, approximately 87.7% (6). Once LM occurs, patients generally lose the opportunity for curative treatment, and the therapeutic strategy must shift towards systemic therapy and comprehensive management (7). However, in clinical practice, some patients may not have clear evidence of LM at the time of initial diagnosis but may have occult metastasis or experience metastatic progression early in treatment (8). Therefore, there is an urgent need to establish an accurate and generalizable predictive model for LM in pancreatic malignancies to identify high-risk patients early, optimize staging assessments, adjust treatment strategies, and improve prognosis.

Currently, LM risk assessment primarily relies on imaging examinations and conventional clinical pathological indicators. However, pancreatic malignancies exhibit a certain degree of histological heterogeneity, with different pathological subtypes potentially displaying varying biological behaviors and metastasis patterns. Traditional methods may have limitations in capturing complex nonlinear relationships and high-dimensional variable interactions. In contrast, machine learning (ML) methods can efficiently integrate multidimensional clinical information and uncover latent patterns, providing new perspectives for the personalized prediction of metastatic risks (9,10). While ML algorithms have been widely applied in risk and prognostic studies for various cancers, including PC, colorectal cancer, and gastric cancer (11-14), there remains a lack of predictive models for LM in pancreatic malignancies that are based on large-scale population data, encompass multiple histological subtypes, and offer robust interpretability.

Based on this, the aim of this study is to construct a predictive model for LM in pancreatic malignancies that covers a broad range of histological types. We focus on the overall population of pancreatic malignancies, including both common types and rare, highly invasive subtypes. Through stratified evaluation systems, we assess the model’s discriminative ability and calibration performance across different histological populations. The goal is to provide a clinically interpretable risk stratification tool that aids in the early identification of patients at high risk for LM and supports decision-making, while further exploring its potential application value in precision medicine. We present this article in accordance with the TRIPOD reporting checklist (available at https://gs.amegroups.com/article/view/10.21037/gs-2026-1-0088/rc).


Methods

Study population

This study extracted data on PC patients from 17 regions covered by the US Surveillance, Epidemiology, and End Results (SEER) database, spanning from 2010 to 2015, with a total of 52,457 cases. The inclusion criteria for patients were as follows: (I) primary pancreatic location; (II) diagnosis between 2010 and 2015; and (III) complete data with no missing values. Exclusion criteria included: (I) patients diagnosed solely by autopsy or death certificate; (II) PC not being the first primary malignancy; and (III) cases with unknown variables. After excluding samples with incomplete data, a total of 12,098 subjects were included for final analysis (Figure 1). A random sampling method was used to divide the subjects into a training set and a validation set at a 7:3 ratio. The training set was used for hyperparameter optimization, while the validation set assessed model performance. The study was conducted in accordance with the Declaration of Helsinki and its subsequent amendments. The study did not require approval from an ethics committee, nor did it require patient consent, as the data was publicly available and devoid of specific personal information.

Figure 1 Flowchart of study design and patient screening. GBM, Gradient Boosting Machine; LASSO, Least Absolute Shrinkage and Selection Operator; LR, Logistic Regression; RF, Random Forest; RFE, Recursive Feature Elimination; SEER, Surveillance, Epidemiology, and End Results; SHAP, SHapley Additive exPlanations; SVM, Support Vector Machine; XGB, XGBoost.

Study variables

This study, based on the SEER database, extracted the clinical and pathological characteristics of each patient. The collected data included the patient’s age, gender, race, marital status, primary site, histological grade, pathological type, tumor size, T stage, N stage, bone metastasis, LM, lung metastasis, brain metastasis, radiotherapy, chemotherapy, and surgical treatment. All PC patients were staged according to the American Joint Committee on Cancer (AJCC) Seventh Edition guidelines.

Data preparation and splitting

To evaluate the generalizability of the predictive model and reduce the risk of overfitting, this study employed a stratified random sampling method (with a fixed random seed of seed =7). The final cleaned dataset was randomly divided into a training set (n=8,469) and a validation set (n=3,629) in a 7:3 ratio. Stratified sampling was used to ensure that the distribution of the outcome variable (synchronous LM) in the training and validation sets was consistent with the original dataset, thus minimizing sampling bias caused by outcome imbalance.

Screening for risk factors and model construction

This study adhered to the principle of “multiple-method statistical validation” in conducting the screening of risk factors for LM in PC. A combination of three feature selection methods—Least Absolute Shrinkage and Selection Operator (LASSO) regression, recursive feature elimination (RFE), and Boruta-Shap—was employed to enhance the reliability and stability of identifying core risk factors. Initially, Spearman correlation analysis was conducted to assess the collinearity between the candidate variables (15). The results showed that no strong correlations were found among the 16 candidate variables, thereby eliminating significant multicollinearity interference for subsequent feature selection. Subsequently, LASSO regression was employed for feature dimensionality reduction and selection. The core advantage of LASSO regression lies in its use of L1 regularization to penalize coefficients, effectively shrinking the coefficients of redundant variables to zero. This simultaneously performs both feature selection and dimensionality reduction (16). It is important to note that for unordered multi-class variables, we applied One-Hot Encoding to ensure their effective inclusion in the LASSO model. RFE performs iterative removal of features with lower contributions, enabling stepwise dimensionality reduction while effectively retaining core features that have a significant contribution to outcome prediction (17). The Boruta-Shap method combines the stability of the Boruta algorithm with the interpretability of SHapley Additive exPlanations (SHAP) values. By comparing the importance of real features with randomly generated “shadow features”, it accurately identifies the features that have a significant impact on the prediction outcomes (18). Finally, we integrated the feature sets selected by the three methods by taking the intersection of the features identified by LASSO regression, RFE, and Boruta-Shap. Only the variables selected by all three methods were retained, forming the final set of core risk factors. This integration strategy effectively eliminates false positive features that may arise from any single method, enhancing the reliability and stability of the core risk factors.

ML model development

Based on the feature selection techniques mentioned above, this study developed six ML models: Logistic Regression (LR) (19), Support Vector Machine (SVM) (20), Random Forest (RF) (21), Gradient Boosting Machine (GBM) (22), XGBoost (XGB) (23), and CatBoost (24). The ML models were evaluated using the following metrics: accuracy, sensitivity (Recall), specificity, F1 score, and area under the receiver operating characteristic (ROC) curve (AUC). Subsequently, precision-recall (PR) curves were constructed to compare the recognition ability of different models, and Calibration Curves were created to assess the calibration performance of the models. decision curve analysis (DCA) is a novel algorithm commonly used to estimate the net benefit value of a model under different thresholds. Based on the best-performing model, SHAP methodology was applied to analyze the contribution of each feature to the predictions.

Statistical analysis

This study primarily used Python (version 3.11.9) and R (version 4.5.1) for data cleaning, statistical analysis, and model construction. In the baseline feature description, continuous variables were expressed as mean ± standard deviation (mean ± SD) if they followed a normal distribution, with group comparisons made using the independent sample Student’s t-test. If the variables exhibited skewed distribution, they were presented as median and interquartile range [median (IQR)], and group comparisons were conducted using the Mann-Whitney U test. Categorical variables were represented as frequency (percentage), with group comparisons made using the Chi-square test or Fisher’s exact test. To rigorously assess the comparability of baseline characteristics between the training and validation sets, the standardized mean difference (SMD) was calculated, with SMD <0.1 used as the criterion for good covariate balance. The DeLong test was employed to statistically compare the differences in AUC values between the candidate ML models. All statistical tests were two-sided, with a P value of <0.05 considered statistically significant.

The ML model construction was based on the scikit-learn (version 1.2.2) framework. Model performance evaluation metrics included the AUC, sensitivity, specificity, and calibration curve, with DCA used to assess clinical net benefit. Model interpretation and feature importance analysis were performed using SHAP (shap, version 0.44.1). Data processing and visualization primarily relied on numpy (version 1.24.4), pandas, and matplotlib. Model saving and deployment were handled using joblib, and an online risk calculator was developed using Streamlit for web deployment.


Results

Clinical and pathological characteristics

Among all the included patients, the dataset was divided into a training set (n=8,469) and a validation set (n=3,629) in a 7:3 ratio. The two groups showed balanced distributions in terms of key clinical and pathological characteristics (Table 1). For instance, the incidence of LM in the training and validation sets was 18.3% and 18.6%, respectively, with no statistically significant difference (P=0.638, SMD =0.009). There were no significant differences between the two groups in terms of age, gender, race, histological grade, primary site, pathological type, T stage, N stage, tumor size, and treatment-related variables (surgery, radiotherapy, chemotherapy). The SMD values for all variables were less than 0.1, indicating good comparability between the training and validation sets.

Table 1

The demographic and clinical characteristics of patients with pancreatic cancer in the training and validation cohorts

Variable Category Training cohort (n=8,469) Validation cohort (n=3,629) SMD χ2 P
Age (years) 64.88±11.88 64.77±11.82 0.009 0.65
Tumor size (mm) 35.00 [25.00, 48.00] 35.00 [25.00, 49.00] 0.007 0.42
Gender 0 = Female 4,042 (47.7) 1,794 (49.4) 0.034 2.968 0.08
1 = Male 4,427 (52.3) 1,835 (50.6)
Race 0 = White 6,735 (79.5) 2,867 (79.0) 0.016 0.706 0.70
1 = Black 921 (10.9) 396 (10.9)
2 = Other 813 (9.6) 366 (10.1)
Marital status 0 = Married 5,291 (62.5) 2,254 (62.1) 0.013 0.429 0.80
1 = Single/unmarried or domestic partner 1,254 (14.8) 531 (14.6)
2 = Separated/divorced/widowed 1,924 (22.7) 844 (23.3)
Primary site 0 = Head of pancreas 5,013 (59.2) 2,169 (59.8) 0.023 2.944 0.40
1 = Body of pancreas 977 (11.5) 446 (12.3)
2 = Tail of pancreas 1,329 (15.7) 541 (14.9)
3 = Other 1,150 (13.6) 473 (13.0)
Histology 0 = Adenocarcinoma, NOS 4,291 (50.7) 1,830 (50.4) 0.025 2.736 0.60
1 = Infiltrating duct carcinoma, NOS 2,013 (23.8) 865 (23.8)
2 = Neuroendocrine carcinoma, NOS 635 (7.5) 249 (6.9)
3 = Carcinoid tumor, NOS 520 (6.1) 242 (6.7)
4 = Other 1,010 (11.9) 443 (12.2)
Grade 0 = I 1,727 (20.4) 725 (20.0) 0.033 3.063 0.38
1 = II 3,457 (40.8) 1,484 (40.9)
2 = III 3,089 (36.5) 1,317 (36.3)
3 = IV 196 (2.3) 103 (2.8)
T stage 0 = T1 705 (8.3) 320 (8.8) 0.026 2.478 0.47
1 = T2 1,668 (19.7) 678 (18.7)
2 = T3 4,888 (57.7) 2,096 (57.8)
3 = T4 1,208 (14.3) 535 (14.7)
N stage 0 = N0 4,191 (49.5) 1,763 (48.6) 0.018 0.833 0.36
1 = N1 4,278 (50.5) 1,866 (51.4)
Surgery 0 = No 3,180 (37.5) 1,343 (37.0) 0.011 0.318 0.57
1 = Yes 5,289 (62.5) 2,286 (63.0)
Radiation 0 = No 6,674 (78.8) 2,837 (78.2) 0.015 0.598 0.43
1 = Yes 1,795 (21.2) 792 (21.8)
Chemotherapy 0 = No 3,403 (40.2) 1,450 (40.0) 0.005 0.054 0.81
1 = Yes 5,066 (59.8) 2,179 (60.0)
Bone metastasis 0 = No 8,347 (98.6) 3,582 (98.7) 0.013 0.390 0.53
1 = Yes 122 (1.4) 47 (1.3)
Brain metastasis 0 = No 8,458 (99.9) 3,627 (99.9) 0.025 1.323 0.25
1 = Yes 11 (0.1) 2 (0.1)
Lung metastasis 0 = No 8,100 (95.6) 3,469 (95.6) 0.003 0.016 0.898
1 = Yes 369 (4.4) 160 (4.4)
Liver metastasis 0 = No 6,922 (81.7) 2,953 (81.4) 0.009 0.221 0.63
1 = Yes 1,547 (18.3) 676 (18.6)

Data are presented as mean ± standard deviation, median (interquartile range) or number (%). N, node; NOS, not otherwise specified; SMD, standardized mean difference; T, tumor.

Variable filtering

In the training set, baseline clinical characteristics of patients with LM were compared to those without LM (Table 2). The results showed significant differences between the two groups in terms of gender, race, primary site, pathological type, histological grade, T stage, N stage, tumor size, surgery, radiotherapy, and other distant metastases (bone metastasis, brain metastasis, lung metastasis) (P<0.05). These findings suggest that these variables may be associated with the occurrence of LM.

Table 2

Baseline characteristics of pancreatic cancer with and without liver metastasis in the training cohort

Variable Category Liver metastasis =0
(n=6,922)
Liver metastasis =1
(n=1,547)
SMD χ2 P
Age (years) 64.99±11.91 64.36±11.71 0.054 0.05
Tumor size (mm) 34.00 [25.00, 45.00] 43.00 [32.00, 58.00] 0.477 <0.001
Gender 0 = Female 3,403 (49.2) 639 (41.3) 0.158 31.282 <0.001
1 = Male 3,519 (50.8) 908 (58.7)
Race 0 = White 5,492 (79.3) 1,243 (80.3) 0.082 9.379 0.009
1 = Black 736 (10.6) 185 (12.0)
2 = Other 694 (10.0) 119 (7.7)
Marital status 0 = Married 4,341 (62.7) 950 (61.4) 0.055 3.914 0.14
1 = Single/unmarried or domestic partner 1,000 (14.4) 254 (16.4)
2 = Separated/divorced/widowed 1,581 (22.8) 343 (22.2)
Primary site 0 = Head of pancreas 4,377 (63.2) 636 (41.1) 0.454 267.781 <0.001
1 = Body of pancreas 737 (10.6) 240 (15.5)
2 = Tail of pancreas 941 (13.6) 388 (25.1)
3 = Other 867 (12.5) 283 (18.3)
Histology 0 = Adenocarcinoma, NOS 3,292 (47.6) 999 (64.6) 0.549 333.442 <0.001
1 = Infiltrating duct carcinoma, NOS 1,899 (27.4) 114 (7.4)
2 = Neuroendocrine carcinoma, NOS 480 (6.9) 155 (10.0)
3 = Carcinoid tumor, NOS 465 (6.7) 55 (3.6)
4 = Other 786 (11.4) 224 (14.5)
Grade 0 = I 1,549 (22.4) 178 (11.5) 0.359 260.515 <0.001
1 = II 2,948 (42.6) 509 (32.9)
2 = III 2,304 (33.3) 785 (50.7)
3 = IV 121 (1.7) 75 (4.8)
T stage 0 = T1 654 (9.4) 51 (3.3) 0.447 454.239 <0.001
1 = T2 1,147 (16.6) 521 (33.7)
2 = T3 4,271 (61.7) 617 (39.9)
3 = T4 850 (12.3) 358 (23.1)
N stage 0 = N0 3,315 (47.9) 876 (56.6) 0.176 38.594 <0.001
1 = N1 3,607 (52.1) 671 (43.4)
Surgery 0 = No 1,842 (26.6) 1,338 (86.5) 1.516 1,933.325 <0.001
1 = Yes 5,080 (73.4) 209 (13.5)
Radiation 0 = No 5,212 (75.3) 1,462 (94.5) 0.557 279.337 <0.001
1 = Yes 1,710 (24.7) 85 (5.5)
Chemotherapy 0 = No 2,769 (40.0) 634 (41.0) 0.020 0.505 0.47
1 = Yes 4,153 (60.0) 913 (59.0)
Bone metastasis 0 = No 6,880 (99.4) 1,467 (94.8) 0.275 185.549 <0.001
1 = Yes 42 (0.6) 80 (5.2)
Brain metastasis 0 = No 6,918 (99.9) 1,540 (99.5) 0.078 15.186 <0.001
1 = Yes 4 (0.1) 7 (0.5)
Lung metastasis 0 = No 6,766 (97.7) 1,334 (86.2) 0.434 402.312 <0.001
1 = Yes 156 (2.3) 213 (13.8)

Data are presented as mean ± standard deviation, median (interquartile range) or number (%). N, node; NOS, not otherwise specified; SMD, standardized mean difference; T, tumor.

To avoid the impact of multicollinearity on model stability, we conducted a Spearman correlation analysis of all candidate variables in the training set and plotted a correlation heatmap (Figure 2). The results showed that most variables had low correlation coefficients (|ρ|<0.7), with no significant multicollinearity observed. Therefore, all variables were included in the subsequent feature selection and modeling analysis. After excluding multicollinearity, feature selection was performed sequentially using LASSO regression, RFE, and Boruta-Shap algorithms.

Figure 2 Heat map of the correlation of features. N, node; T, tumor.

To further identify key features associated with LM in PC, we used LASSO regression for feature selection in the training set. The LASSO coefficient path plot showed that as the penalty parameter λ increased, the regression coefficients of some variables gradually shrank and approached zero (Figure 3A). Subsequently, the optimal λ value was determined through 10-fold cross-validation (Figure 3B), and two candidate models were provided, λ.min and λ.1se. At the penalty strength corresponding to λ.1se, 11 non-zero coefficient variables were retained (Figure 3C). Among these, bone metastasis, lung metastasis, histological grade, and tumor size were positively associated with the occurrence of LM, while surgery, radiotherapy, and age had negative regression coefficients, suggesting their association with a lower risk of LM. As a comparison, under λ.min, the model included 16 non-zero coefficient variables (Figure 3D), indicating that when the penalty is relatively weaker, more variables are included in the model. After considering the model’s stability and clinical interpretability, this study ultimately selected the feature set corresponding to λ.1se as the primary analysis result.

Figure 3 Selection of predictive variables using LASSO logistic regression. (A) Coefficient profiles of candidate variables along the LASSO regularization path. Each colored line depicts the variation trend of a variable’s coefficient along the regularization path with increasing λ values. The red dashed line indicates λ.min, while the gray dotted line represents λ.1se. (B) Ten-fold cross-validation curve for identifying the optimal penalty parameter λ. The minimum binomial deviance corresponds to λ.min, and the more parsimonious parameter λ.1se was selected in accordance with the one-standard-error rule. (C) Coefficients of the features retained at λ.1se (under stronger penalization, with 11 variables retained). (D) Coefficients of features retained at λ.min (less penalized, 16 variables retained). Positive coefficients indicate risk factors for pancreatic cancer liver metastasis, whereas negative coefficients denote protective factors against this metastatic process. LASSO, Least Absolute Shrinkage and Selection Operator; N, node; T, tumor.

In the RFE analysis, we used XGBoost as the base model to iteratively rank and prune the features (Figure 4A). As the number of features included increased, the model performance showed a rapid rise followed by gradual stabilization. When the number of retained features reached 15, the average AUC from 10-fold cross-validation peaked at 0.888. These 15 features identified as key candidate variables included: race, age, histological grade, tumor size, T stage, N stage, surgery, chemotherapy, radiotherapy, bone metastasis, lung metastasis, pathological type, primary site, gender, and marital status.

Figure 4 Correlation of feature variables and feature selection process in machine learning. (A) XGBoost-RFE feature selection results. Cross-validation AUC under optimal hyperparameter combination. (B) Significant features identified by Boruta-Shap algorithm for liver metastasis prediction in pancreatic cancer. (C) Venn diagrams compare the variables selected by the three different methods, showing the overlap of the selected variables. AUC, area under the curve; LASSO, Least Absolute Shrinkage and Selection Operator; N, node; RFE, recursive feature elimination; T, tumor.

Finally, we applied the XGBoost-based Boruta-Shap algorithm. This algorithm utilizes the SHAP method to calculate the importance scores of each variable and determines variable importance by iteratively comparing with randomly generated “shadow features”. As shown in Figure 4B, all candidate variables were ranked according to the average SHAP Z-score. Variables with significantly higher scores than the shadow features were classified as “Accepted”, while those with lower scores were “Rejected”. A total of 13 variables were identified as key predictors influencing the occurrence of LM in PC, including race, age, histological grade, tumor size, T stage, N stage, surgery, chemotherapy, radiotherapy, bone metastasis, lung metastasis, pathological type, and primary site.

Based on the results from the three feature selection methods, the final core predictive features were determined to be surgery, radiotherapy, tumor size, age, histological grade, primary site, pathological type, T stage, lung metastasis, and bone metastasis (Figure 4C). These core features were used for the subsequent construction and validation of the ML models.

Predictive performance of ML models

Table 3 summarizes the performance metrics of the six ML models in both the training and validation sets. Overall, all models demonstrated excellent predictive performance in the training set. Among them, the CatBoost model achieved the highest AUC value [0.922, 95% confidence interval (CI): 0.915–0.928], followed by Random Forest (0.920, 95% CI: 0.913–0.927) and XGBoost (0.907, 95% CI: 0.899–0.914) (Figure 5). However, to ensure the clinical applicability and predictive reliability of the models, the final selection of the best model was primarily based on its comprehensive performance in the internal validation set, incorporating both quantitative metrics and curve evaluation evidence for multidimensional consideration.

Table 3

Performance of various machine learning models for liver metastasis in pancreatic cancer

Metrics LR SVM RF GBM XGB CatBoost
Cutoff 0.512 0.138 0.421 0.176 0.404 0.410
   Training AUC 0.882 (0.873–0.890) 0.897 (0.888–0.905) 0.920 (0.913–0.927) 0.895 (0.887–0.903) 0.907 (0.899–0.914) 0.922 (0.915–0.928)
   Testing AUC 0.871 (0.858–0.884) 0.869 (0.855–0.882) 0.877 (0.863–0.890) 0.875 (0.862–0.889) 0.879 (0.865–0.892) 0.879 (0.867–0.892)
Accuracy 0.792 0.785 0.792 0.790 0.784 0.788
   Sensitivity 0.853 (0.824–0.879) 0.866 (0.838–0.891) 0.857 (0.828–0.883) 0.851 (0.822–0.878) 0.874 (0.846–0.898) 0.869 (0.841–0.894)
   Specificity 0.778 (0.763–0.793) 0.766 (0.751–0.782) 0.777 (0.762–0.792) 0.776 (0.761–0.791) 0.764 (0.748–0.779) 0.770 (0.755–0.785)
   PPV 0.464 (0.435–0.492) 0.455 (0.427–0.483) 0.464 (0.436–0.492) 0.461 (0.433–0.489) 0.454 (0.427–0.482) 0.460 (0.432–0.488)
   NPV 0.959 (0.951–0.967) 0.962 (0.954–0.970) 0.960 (0.952–0.968) 0.959 (0.950–0.966) 0.964 (0.956–0.971) 0.963 (0.955–0.970)
F1 score 0.601 0.596 0.602 0.598 0.598 0.601
MCC 0.517 0.514 0.519 0.513 0.517 0.520
Kappa 0.476 0.468 0.478 0.472 0.470 0.475
PLR 3.845 3.708 3.848 3.803 3.703 3.781
NLR 0.189 0.174 0.184 0.192 0.165 0.170
mAP 0.552 0.533 0.575 0.575 0.584 0.585

Data are presented as point estimates or point estimates with 95% confidence intervals. AUC, area under the curve; GBM, Gradient Boosting Machine; LR, Logistic Regression; mAP, mean average precision; MCC, Matthews correlation coefficient; NLR, negative likelihood ratio; NPV, negative predictive value; PLR, positive likelihood ratio; PPV, positive predictive value; RF, Random Forest; SVM, Support Vector Machine; XGB, XGBoost.

Figure 5 The ROC curve for evaluating the diagnostic performance of machine learning models. (A) The training set ROC curve, (B) the validation set ROC curve. AUC, area under the curve; CI, confidence interval; ROC, receiver operating characteristic; SVM, Support Vector Machine.

In the independent validation set, quantitative evaluation metrics showed that all six models maintained robust predictive performance, with AUC values ranging from 0.869 (SVM) to 0.879 (XGBoost and CatBoost). The GBM model demonstrated highly competitive predictive performance with an AUC of 0.875 (95% CI: 0.862–0.889). Although the DeLong test indicated that XGBoost and CatBoost achieved statistically higher AUCs than the GBM (P=0.0002 and P=0.0006, respectively; see Table S1), the absolute differences were negligible (ranging from 0.003 to 0.005). Importantly, considering that AUC primarily measures discriminative rank rather than clinical utility, three key curve analyses consistently confirmed the comprehensive superiority of the GBM model in terms of calibration accuracy and net clinical benefit:

First, the PR curve analysis (Figure 6A) showed that the GBM model achieved an average precision (AP) of 0.575. This value not only significantly outperformed the other three models but was also comparable to the best-performing XGBoost (0.584) and CatBoost (0.585). This indicates that, even in the context of imbalanced data distribution for LM events, the GBM model still exhibits strong robustness in accurately identifying positive patients.

Figure 6 Performance comparison of different models. (A) PR curve analysis, (B) calibration curve analysis, (C) DCA curves. DCA, decision curve analysis; PR, precision-recall; SVM, Support Vector Machine.

Next, the calibration curve (Figure 6B) demonstrated that the GBM model had the best alignment with the ideal prediction curve. This indicates that the predicted probabilities of LM output by the model closely match the actual occurrence probabilities, reflecting its excellent calibration ability. This characteristic is crucial for clinical decision-making, as it directly ensures the authenticity and reliability of the risk assessment values provided to clinicians.

Finally, the DCA (Figure 6C) further established the clinical utility of the GBM model. Compared to the other five models, the GBM model consistently generated higher Net Clinical Benefit (NCB) across a wider range of threshold probabilities. This means that when applied to guide clinical intervention strategies for PC patients, the model effectively balances the “benefit of identifying high-risk patients” with the “potential harm from overtreatment,” demonstrating exceptional value in clinical decision support.

Model performance validation across histological subgroups

To assess the model’s generalizability across different histological subtypes, we conducted a stratified evaluation based on histological origin in the validation set. The main subgroups included the pancreatic ductal adenocarcinoma (PDAC) group (comprising adenocarcinoma and invasive ductal carcinoma, n=2,695) and the pancreatic neuroendocrine neoplasms (PanNENs) group (comprising neuroendocrine carcinoma and carcinoid, n=491). Given the complex composition of histological types, the “Other” group (accounting for 12.2% of the validation cohort) was included in the overall model performance assessment but was not analyzed separately for subgroup visualization. The results (Figure 7) showed that the GBM model maintained good discriminative ability in both subgroups. Specifically, the model achieved an AUC of 0.861 in the PDAC group and an AUC of 0.887 in the PanNENs group (Figure 7A). This suggests that the model performed well in the PDAC-dominant population while maintaining its risk identification capability for the relatively rare subtype (PanNENs). Furthermore, the subgroup calibration curves (Figure 7B) indicated that the predicted probabilities aligned well with the actual observed probabilities in both subgroups, suggesting that the model exhibits good calibration performance across different histological backgrounds.

Figure 7 Model performance analysis plot for histological subtypes of pancreatic tumors. (A) ROC curves by histology subgroup, (B) calibration by histology subgroup. AUC, area under the curve; NEC, neuroendocrine carcinoma; PanNENs, pancreatic neuroendocrine neoplasms; PDAC, pancreatic ductal adenocarcinoma; ROC, receiver operating characteristic.

In summary, although models such as XGBoost and CatBoost demonstrated comparable quantitative predictive performance to the GBM model in the validation set, the GBM model showed a more decisive overall advantage across three core evaluation dimensions: the PR curve, calibration curve, and DCA decision curve. Given that these curve-based evaluations more objectively reflect the model’s clinical applicability and predictive robustness in real-world scenarios, the GBM model was ultimately established as the optimal model for predicting synchronous LM in PC patients in this study.

Feature importance in ML models

This study applied the SHAP algorithm to interpret the GBM model, providing traceable clinical insights at both the global (feature-level) and local (individual-level) interpretations.

At the global interpretation level, we first quantified the average contribution of each feature to the model’s predictions. As shown in Figure 8A and Figure 8B, the feature importance ranking based on the mean absolute SHAP values (mean |SHAP|) revealed that the primary tumor surgical status was the variable with the highest contribution, followed by radiotherapy, tumor size, age, histological grade, primary site, pathological type, T stage, lung metastasis, and bone metastasis. These features collectively form the key informational structure used by the model to assess LM risk.

Figure 8 Global model explanation by the SHAP method. (A) SHAP value for subgroup of each variable; (B) bar chart of mean absolute SHAP values ranked by feature importance; (C) the SHAP heatmap clusters hierarchically based on SHAP values; (D) SHAP decision plot. Each line corresponds to a single sample; the spread of lines reflects the variability in feature impact, with the x-axis representing the model output value; (E) SHAP scatter plot (age). This plot shows the relationship between age (x-axis) and its SHAP value (y-axis). The color gradient denotes the Age value (blue = low, red = high); a positive SHAP value indicates that the feature is associated with higher risk (per the plot annotation); (F) SHAP scatter plot (tumor size). This plot displays the relationship between tumor size (x-axis) and its SHAP value (y-axis). The color gradient represents tumor size (blue = small, red = large); a positive SHAP value corresponds to higher risk (per the plot annotation). SHAP, SHapley Additive exPlanations; T, tumor.

Further, the SHAP heatmap (Figure 8C), combined with cluster analysis, visualized the SHAP value patterns at the individual level, presenting the risk feature combinations across different patients. The red areas correspond to higher positive contributions (indicating an increased risk of LM), aggregating patients with high-risk feature profiles. In contrast, the blue areas represent negative contributions (indicating a lower risk), which are more common among low-risk patients or those without LM.

Meanwhile, the SHAP decision plot (Figure 8D) illustrates the cumulative effect of each feature on the prediction, starting from the baseline value. It further highlights how key variables, such as surgical status, radiotherapy, and tumor size, drive the model’s output toward higher or lower risk directions.

Finally, to depict the potential nonlinear relationships between continuous variables and risk, we plotted the SHAP dependence plots (Figure 8E,8F), which display the marginal effects of age and tumor size on the model’s output, as well as their threshold characteristics.

At the local interpretation level, we used SHAP waterfall plots and force plots to reveal the decision-making mechanism of the model at the individual level. Specifically, the SHAP values for each feature reflect the direction and intensity of its contribution to the prediction outcome: red bars (or red regions) indicate that the feature pushes the model output toward the “LM” direction, while blue bars (or blue regions) suggest that the feature pulls the model output toward the “no LM/low risk” direction. In other words, the force plot visually demonstrates how the positive (red) and negative (blue) contributions of each feature collectively drive the model output away from the baseline probability, leading to individualized risk assessment.

As shown in Figure 9A and Figure 9B, the feature combinations of low-risk LM patients primarily have negative contributions, which lower the overall predicted probability. In contrast, Figure 9C and Figure 9D display high-risk LM patients, where multiple key variables have a positive cumulative effect, significantly increasing the model output. These visualizations clearly illustrate how different clinical variable combinations interact at the individual level to influence the model’s final decision and risk stratification.

Figure 9 Local model explanation by the SHAP method. (A,B) SHAP waterfall plot and force plot for pancreatic cancer patients without liver metastasis; (C,D) SHAP waterfall plot and force plot for pancreatic cancer patients with liver metastasis. SHAP, SHapley Additive exPlanations; T, tumor.

Model deployment

To facilitate the use of this predictive model by researchers and clinicians, we developed a user-friendly web application. The application receives clinical feature inputs for new samples through an intuitive web interface (Figure 10) and provides real-time predictions of LM risk in PC patients. The web calculator can be accessed at the following link: https://7x9vgkcao9scjbypfzsdak.streamlit.app/.

Figure 10 The web-based calculator for predicting liver metastasis risk probability in pancreatic cancer. SHAP, SHapley Additive exPlanations.

Discussion

Due to its unique anatomical structure (portal vein return) and immunosuppressive microenvironment, the liver is the most common distant metastatic target organ for pancreatic malignancies (25). The occurrence of LM often signifies disease progression, significantly limiting the possibility of curative surgery and frequently leading to poor survival outcomes (26). In response to this clinical challenge, this study successfully developed and validated a GBM-based ML model, using large-scale population cohort data, to quantitatively assess the risk of LM in pancreatic malignancy patients. The model effectively integrates multidimensional clinical and pathological features, demonstrating excellent predictive performance (AUC =0.875) and good calibration. Notably, by validating the model across different histological subtypes, we further confirmed its robustness, overcoming the prediction challenges posed by the biological heterogeneity of pancreatic tumors. Furthermore, we transformed this complex algorithm into a web-based interactive risk calculator, greatly facilitating its application and translation into clinical risk stratification. A key finding in our model comparison was that the algorithm with the highest AUC was not necessarily the most optimal for clinical use. While complex ensemble methods like CatBoost showed a statistically significant edge in AUC according to the DeLong test (P<0.05), we prioritized the GBM model for its balanced performance. In clinical practice, a model’s calibration—the accuracy of the risk probability it assigns to an individual patient—is often more vital for personalized decision-making than its ability to rank patients (AUC). The GBM’s excellent calibration and superior net benefit in DCA ensure that it provides more reliable and actionable guidance for clinicians.

Previous studies on predicting LM in pancreatic malignancies have often focused on constructing nomograms based on Logistic models (27,28). Although the application of ML in this field has gradually increased in recent years, existing research can generally be divided into two categories: one emphasizes the identification of occult LM based on imaging radiomics (29), while the other focuses on building scalable ML models using conventional clinical variables and performing external validation (30). Current models exhibit different characteristics in terms of imaging dependence, external generalizability, and clinical applicability. Further validation using large-scale population data, while simultaneously considering interpretability, stratified evaluation across different pathological subtypes, and convenient deployment, will help enhance the clinical utility of the model. Additionally, a more systematic feature selection strategy may improve the model’s robustness and interpretability.

In this study, a multi-strategy combined feature selection process was employed during the feature engineering phase, ultimately identifying a set of stable and clinically meaningful predictive features. SHAP interpretability analysis based on this feature set revealed that the surgery of the primary site was of significant importance in predicting LM risk, with its contribution ranking among the highest in the model. This result may partially reflect the association between surgical status and overall patient prognosis and treatment opportunities (31). Previous studies have shown that surgical resection of the primary tumor or LM can effectively delay the progression of LM and improve patient survival outcomes (32). Therefore, surgical status can be considered a key variable in predicting LM risk, reflecting the extent of tumor progression and treatment opportunities for patients (33). However, this finding must be interpreted with caution due to potential reverse causality. In clinical practice, the decision to perform radical surgery is fundamentally contingent upon the pre-existing metastatic status; patients with identified LM are significantly less likely to undergo resection compared to those without. Our data confirms this clinical reality, as 86.5% of patients in the LM group did not receive surgery. Therefore, rather than being a traditional causative risk factor, the high importance of the ‘Surgery’ variable functions as a surrogate indicator of disease severity and resectability. We have included this variable to reflect the real-world clinical profile of patients, but we emphasize that ‘No Surgery’ should be clinically interpreted as a marker for advanced disease stages often associated with synchronous metastasis. Additionally, variables such as radiotherapy, tumor size, and patient age also demonstrated independent predictive value, providing a solid foundation for the model’s accurate predictions.

The core challenge in pancreatic malignancy prediction modeling lies in the significant biological heterogeneity between exocrine tumors (PDAC) and neuroendocrine tumors (PanNENs). PDAC typically presents as highly aggressive with a poor prognosis (34), while PanNENs exhibit unique metastasis patterns and relatively longer survival rates (35,36). Traditional single, general-purpose models often produce predictive biases by overlooking these categorical differences. To overcome this limitation, we incorporated “histological type” as a key predictive factor and conducted rigorous subgroup performance validation. The results showed that, although the model was trained on a mixed cohort, it maintained excellent discriminative ability in both the PDAC subgroup (AUC =0.861) and the PanNENs subgroup (AUC =0.887). This finding confirms that the GBM algorithm successfully learned the complex nonlinear interactions between histological type and other risk factors, allowing it to dynamically adjust the risk score weights accordingly. This robust generalization ability holds significant clinical translational value, especially for patients whose preoperative pathological subtypes are unclear, as the model provides a unified and reliable risk stratification basis.

For the two continuous variables, age and tumor size, we used SHAP dependence plots to examine their relationship with LM. The SHAP dependence plot for age revealed a negative correlation with LM risk, indicating that younger patients are at higher risk for LM compared to older patients. With 65 years old as the threshold where the SHAP value is 0, this suggests that patients over 65 have a lower risk of LM. This finding is consistent with several epidemiological studies, which show that although older age is typically associated with multiple comorbidities and poorer overall survival, early-onset PC often exhibits a more aggressive biological phenotype with a higher propensity for hematogenous metastasis, while older patients may experience a slower disease course (37-39). This finding suggests that in clinical practice, the occult LM in younger patients might be overlooked, and therefore, a heightened level of vigilance is necessary. This population could benefit from more aggressive initial staging and screening. Regarding tumor size, the SHAP dependence plot showed a positive correlation with LM risk. When 36 mm was used as the threshold, patients with tumors larger than 36 mm had a significantly higher risk of LM. This finding is consistent with previous studies, which reported that larger primary tumor size is significantly associated with the occurrence of LM. A larger tumor diameter is considered one of the independent risk factors for LM (40,41).

This study has certain limitations. First, due to the limitations of the SEER database, some imaging features and laboratory indicators with predictive value were not included in the analysis. Future research will consider incorporating specific imaging features and laboratory data such as CA19-9 to further improve the model’s predictive capabilities. Second, although internal validation demonstrated that the model has good robustness and high predictive ability, this study lacks an independent multicenter external validation cohort, which may limit the model’s generalizability across different populations and healthcare settings. Thirdly, the inclusion of concurrent distant metastasis (lung and bone) as predictors might raise concerns regarding temporal data leakage. While these factors are biologically relevant and captured during the same diagnostic window, future studies should focus on developing models that rely solely on preoperative laboratory markers or primary tumor radiomics to predict LM even earlier in the disease course. To further enhance the model’s external generalizability, future research should integrate more independent validation datasets and assess its application in different regions, pathological subtypes, and clinical practices.


Conclusions

We successfully developed and validated a high-performance ML model to accurately predict the risk of LM in patients with PC. Among the six candidate algorithms, the GBM model exhibited the optimal accuracy, robust predictive performance and favorable interpretability. Moreover, the integration of a user-friendly web-based risk calculator further enhances its clinical applicability, enabling personalized risk stratification to support clinical treatment decision-making and rational follow-up planning for PC patients in clinical practice.


Acknowledgments

None.


Footnote

Reporting Checklist: The authors have completed the TRIPOD reporting checklist. Available at https://gs.amegroups.com/article/view/10.21037/gs-2026-1-0088/rc

Peer Review File: Available at https://gs.amegroups.com/article/view/10.21037/gs-2026-1-0088/prf

Funding: This work was supported by Tianjin Key Medical Discipline Construction Project (No. TJYXZDXK-3-009B) and the Tianjin Health Research Project (No. TJWJ2023XK014).

Conflicts of Interest: All authors have completed the ICMJE uniform disclosure form (available at https://gs.amegroups.com/article/view/10.21037/gs-2026-1-0088/coif). The authors have no conflicts of interest to declare.

Ethical Statement: The authors are accountable for all aspects of the work in ensuring that questions related to the accuracy or integrity of any part of the work are appropriately investigated and resolved. The study was conducted in accordance with the Declaration of Helsinki and its subsequent amendments.

Open Access Statement: This is an Open Access article distributed in accordance with the Creative Commons Attribution-NonCommercial-NoDerivs 4.0 International License (CC BY-NC-ND 4.0), which permits the non-commercial replication and distribution of the article with the strict proviso that no changes or edits are made and the original work is properly cited (including links to both the formal publication through the relevant DOI and the license). See: https://creativecommons.org/licenses/by-nc-nd/4.0/.


References

  1. Bray F, Laversanne M, Sung H, et al. Global cancer statistics 2022: GLOBOCAN estimates of incidence and mortality worldwide for 36 cancers in 185 countries. CA Cancer J Clin 2024;74:229-63. [Crossref] [PubMed]
  2. Siegel RL, Kratzer TB, Giaquinto AN, et al. Cancer statistics, 2025. CA Cancer J Clin 2025;75:10-45. [Crossref] [PubMed]
  3. Halbrook CJ, Lyssiotis CA, Pasca di Magliano M, et al. Pancreatic cancer: Advances and challenges. Cell 2023;186:1729-54. [Crossref] [PubMed]
  4. Park W, Chawla A, O'Reilly EM. Pancreatic Cancer: A Review. JAMA 2021;326:851-62. [Crossref] [PubMed]
  5. Ilic M, Ilic I. Epidemiology of pancreatic cancer. World J Gastroenterol 2016;22:9694-705. [Crossref] [PubMed]
  6. Liu Q, Zhang R, Michalski CW, et al. Surgery for synchronous and metachronous single-organ metastasis of pancreatic cancer: a SEER database analysis and systematic literature review. Sci Rep 2020;10:4444. [Crossref] [PubMed]
  7. Mannucci A, Goel A. Advances in pancreatic cancer early diagnosis, prevention, and treatment: The past, the present, and the future. CA Cancer J Clin 2026;76:e70035. [Crossref] [PubMed]
  8. Walma M, Maggino L, Smits FJ, et al. The Difficulty of Detecting Occult Metastases in Patients with Potentially Resectable Pancreatic Cancer: Development and External Validation of a Preoperative Prediction Model. J Clin Med 2024;13:1679. [Crossref] [PubMed]
  9. Yaqoob A, Musheer Aziz R, verma NK. Applications and Techniques of Machine Learning in Cancer Classification: A Systematic Review. Hum-Cent Intell 2023;Syst 3:588-615.
  10. Talebi A, Celis-Morales CA, Borumandnia N, et al. Predicting metastasis in gastric cancer patients: machine learning-based approaches. Sci Rep 2023;13:4163. [Crossref] [PubMed]
  11. Zhou C, Wang Y, Ji MH, et al. Predicting Peritoneal Metastasis of Gastric Cancer Patients Based on Machine Learning. Cancer Control 2020;27:1073274820968900. [Crossref] [PubMed]
  12. Zhang F, Wu G, Chen N, et al. The predictive value of radiomics-based machine learning for peritoneal metastasis in gastric cancer patients: a systematic review and meta-analysis. Front Oncol 2023;13:1196053. [Crossref] [PubMed]
  13. Guo Z, Zhang Z, Liu L, et al. Explainable machine learning for predicting lung metastasis of colorectal cancer. Sci Rep 2025;15:13611. [Crossref] [PubMed]
  14. Chen W, Zhou B, Jeon CY, et al. Machine learning versus regression for prediction of sporadic pancreatic cancer. Pancreatology 2023;23:396-402. [Crossref] [PubMed]
  15. Liu R, Wu S, Kuang T, et al. Prediction system for gallbladder cancer distant metastasis a retrospective cohort study based on machine learning. Int J Surg 2025;111:7906-20. [Crossref] [PubMed]
  16. Li X, Jacobucci R. Regularized structural equation modeling with stability selection. Psychol Methods 2022;27:497-518. [Crossref] [PubMed]
  17. Yang M, Wu Y, Yang XB, et al. Establishing a prediction model of severe acute mountain sickness using machine learning of support vector machine recursive feature elimination. Sci Rep 2023;13:4633. [Crossref] [PubMed]
  18. Samara MN, Harry KD. Integrating Boruta, LASSO, and SHAP for Clinically Interpretable Glioma Classification Using Machine Learning. BioMedInformatics 2025;5:34.
  19. Pan X, Xu Y. A Safe Feature Elimination Rule for L(1)-Regularized Logistic Regression. IEEE Trans Pattern Anal Mach Intell 2022;44:4544-54. [Crossref] [PubMed]
  20. Rezvani S, Wu J. Handling Multi-Class Problem by Intuitionistic Fuzzy Twin Support Vector Machines Based on Relative Density Information. IEEE Trans Pattern Anal Mach Intell 2023;45:14653-64. [Crossref] [PubMed]
  21. Sarica A, Cerasa A, Quattrone A. Random Forest Algorithm for the Classification of Neuroimaging Data in Alzheimer's Disease: A Systematic Review. Front Aging Neurosci 2017;9:329. [Crossref] [PubMed]
  22. Branstad-Spates EH, Castano-Duque L, Mosher GA, et al. Gradient boosting machine learning model to predict aflatoxins in Iowa corn. Front Microbiol 2023;14:1248772. [Crossref] [PubMed]
  23. Chen T, Guestrin C. XGBoost: A Scalable Tree Boosting System. In KDD '16: Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining. Association for Computing Machinery, New York, NY, USA, 2016:785-94.
  24. Prokhorenkova L, Gusev G, Vorobev A, et al. CatBoost: unbiased boosting with categorical features. In: Advances in Neural Information Processing Systems. Curran Associates, Inc. 2018. Available online: https://papers.nips.cc/paper_files/paper/2018/hash/14491b756b3a51daac41c24863285549-Abstract.html
  25. Fabian A, Stegner S, Miarka L, et al. Metastasis of pancreatic cancer: An uninflamed liver micromilieu controls cell growth and cancer stem cell properties by oxidative phosphorylation in pancreatic ductal epithelial cells. Cancer Lett 2019;453:95-106. [Crossref] [PubMed]
  26. Agarwal T, Mohanraj K, Pai A. Prolonged survival in metastatic pancreatic cancer: a case of multimodal therapy. Clinical Medicine 2025;25:100449.
  27. He C, Zhong L, Zhang Y, et al. Development and validation of a nomogram to predict liver metastasis in patients with pancreatic ductal adenocarcinoma: a large cohort study. Cancer Manag Res 2019;11:3981-91. [Crossref] [PubMed]
  28. Chen S, Chen S, Lian G, et al. Development and validation of a novel nomogram for pretreatment prediction of liver metastasis in pancreatic cancer. Cancer Med 2020;9:2971-80. [Crossref] [PubMed]
  29. Yuan Z, Shu Z, Peng J, et al. Prediction of postoperative liver metastasis in pancreatic ductal adenocarcinoma based on multiparametric magnetic resonance radiomics combined with serological markers: a cohort study of machine learning. Abdom Radiol (NY) 2024;49:117-30. [Crossref] [PubMed]
  30. Zhu H, Zhou Y, Shen D, et al. An interpretable machine learning model for predicting early liver metastasis after pancreatic cancer surgery. BMC Cancer 2025;25:1117. [Crossref] [PubMed]
  31. Nagai M, Wright MJ, Ding D, et al. Oncologic resection of pancreatic cancer with isolated liver metastasis: Favorable outcomes in select patients. J Hepatobiliary Pancreat Sci 2023;30:1025-35. [Crossref] [PubMed]
  32. Zhao P, Wang Z, Xue K, et al. Synchronous surgery combined preoperative chemotherapy benefits patients suffering pancreatic ductal adenocarcinoma with liver metastases: a systematic review and meta-analysis. Sci Rep 2025;15:28403. [Crossref] [PubMed]
  33. Clements N, Gaskins J, Martin RCG 2nd. Surgical Outcomes in Stage IV Pancreatic Cancer with Liver Metastasis Current Evidence and Future Directions: A Systematic Review and Meta-Analysis of Surgical Resection. Cancers (Basel) 2025;17:688. [Crossref] [PubMed]
  34. Schmidtlein PM, Volz C, Hackel A, et al. Activation of a Ductal-to-Endocrine Transdifferentiation Transcriptional Program in the Pancreatic Cancer Cell Line PANC-1 Is Controlled by RAC1 and RAC1b through Antagonistic Regulation of Stemness Factors. Cancers (Basel) 2021;13:5541. [Crossref] [PubMed]
  35. Dasari A, Shen C, Halperin D, et al. Trends in the Incidence, Prevalence, and Survival Outcomes in Patients With Neuroendocrine Tumors in the United States. JAMA Oncol 2017;3:1335-42. [Crossref] [PubMed]
  36. Cives M, Strosberg JR. Gastroenteropancreatic Neuroendocrine Tumors. CA Cancer J Clin 2018;68:471-87. [Crossref] [PubMed]
  37. Li Z, Zhang X, Sun C, et al. Global, regional, and national burdens of early onset pancreatic cancer in adolescents and adults aged 15-49 years from 1990 to 2019 based on the Global Burden of Disease Study 2019: a cross-sectional study. Int J Surg 2024;110:1929-40. [Crossref] [PubMed]
  38. Ansari D, Althini C, Ohlsson H, et al. Early-onset pancreatic cancer: a population-based study using the SEER registry. Langenbecks Arch Surg 2019;404:565-71. [Crossref] [PubMed]
  39. Primavesi F, Stättner S, Schlick K, et al. Pancreatic cancer in young adults: changes, challenges, and solutions. Onco Targets Ther 2019;12:3387-400. [Crossref] [PubMed]
  40. S D. Risk factors of liver metastasis from advanced pancreatic adenocarcinoma: a large multicenter cohort study. World J Surg Oncol 2017;15:120. [Crossref] [PubMed]
  41. Cao BY, Tong F, Zhang LT, et al. Risk factors, prognostic predictors, and nomograms for pancreatic cancer patients with initially diagnosed synchronous liver metastasis. World J Gastrointest Oncol 2023;15:128-42. [Crossref] [PubMed]
Cite this article as: Chen Z, Song P, Fang H, Tian W, Zhao J, Liu Y, Wang C, Zhao Y, Cao L. Development and validation of a machine learning model for predicting liver metastasis in pancreatic cancer: a retrospective cohort study. Gland Surg 2026;15(4):85. doi: 10.21037/gs-2026-1-0088

Download Citation