Development and validation of an explainable web-based machine learning model for predicting early postoperative hypocalcemia in patients with differentiated thyroid cancer
Original Article

Development and validation of an explainable web-based machine learning model for predicting early postoperative hypocalcemia in patients with differentiated thyroid cancer

Jiaqi Ye1, Yuanyuan Wang2 ORCID logo

1Department of Operating Room, The First Hospital of China Medical University, Shenyang, China; 2Department of Vascular and Thyroid Surgery, The First Hospital of China Medical University, Shenyang, China

Contributions: (I) Conceptualization and design: J Ye; (II) Administrative support: Y Wang; (III) Provision of study materials or patients: J Ye; (IV) Collection and assembly of data: Both authors; (V) Data analysis and interpretation: Both authors; (VI) Manuscript writing: Both authors; (VII) Final Approval of manuscript: Both authors.

Correspondence to: Yuanyuan Wang, BD. Department of Vascular and Thyroid Surgery, The First Hospital of China Medical University, No. 155 Nanjing North Street, Heping District, Shenyang 110001, China. Email: wangyy@cmu1h.com.

Background: Early postoperative hypocalcemia is one of the most common complications after surgery for differentiated thyroid cancer (DTC) and may lead to prolonged hospitalization and impaired postoperative recovery. Most existing prediction tools are based on conventional regression methods and may be insufficient to capture complex interactions among clinical variables. This study aimed to develop and validate an interpretable machine learning model and a web-based calculator for early prediction of postoperative hypocalcemia in patients with DTC.

Methods: This retrospective cohort study included 1,222 patients with DTC who underwent surgery at a tertiary hospital between February 2021 and October 2025. Candidate predictors were screened using a combined feature selection strategy incorporating least absolute shrinkage and selection operator (LASSO) regression and the Boruta algorithm. Eight machine learning algorithms were developed and compared. Model performance was evaluated using the area under the receiver operating characteristic curve (AUC) and calibration metrics. Model interpretability was assessed using SHapley Additive exPlanations (SHAP), and the optimal model was deployed as a web-based calculator.

Results: Seven consensus variables were selected for final model construction. Among the eight algorithms, the light gradient boosting machine (LGBM) achieved the best overall performance in the testing cohort, with an AUC of 0.875 [95% confidence interval (CI): 0.838–0.912] and a Brier score of 0.132, indicating good discrimination and calibration. SHAP analysis showed that bilateral central lymph node dissection (CLND), severe preoperative vitamin D deficiency, and thyroid capsular invasion were the most influential predictors of postoperative hypocalcemia. The final model was implemented as a freely accessible web-based clinical tool.

Conclusions: We developed an explainable web-based LGBM model that showed good predictive performance for early postoperative hypocalcemia in patients with DTC. By combining accurate risk estimation with transparent interpretation, this tool may help clinicians identify high-risk patients and support individualized perioperative management.

Keywords: Differentiated thyroid carcinoma; postoperative hypocalcemia; machine learning (ML); explainable prediction model; web-based calculator


Submitted Apr 10, 2026. Accepted for publication Jun 12, 2026. Published online Jul 23, 2026.

doi: 10.21037/gs-2026-0211


Highlight box

Key findings

• We developed and validated an explainable machine learning model for the early prediction of postoperative hypocalcemia in patients with differentiated thyroid cancer.

• The light gradient boosting machine (LGBM) model demonstrated the best overall performance, with good discrimination and calibration.

• SHapley Additive exPlanations (SHAP) analysis identified key predictors, including severe preoperative vitamin D deficiency, bilateral central lymph node dissection (CLND), and thyroid capsular invasion.

• A web-based calculator was developed to facilitate clinical application.

What is known and what is new?

• Postoperative hypocalcemia is a frequent and burdensome complication of thyroidectomy. Existing predictive tools predominantly rely on traditional logistic regression, which struggles to capture complex, non-linear interactions between surgical extent and preoperative metabolic status.

• This study demonstrates the superiority of the LGBM model in modeling the non-linear physiological pathways contributing to hypocalcemia. By integrating SHAP values and deploying the model as a freely accessible web-based calculator, this framework effectively bridges algorithmic complexity with clinical transparency.

What is the implication, and what should change now?

• Surgeons may utilize this web-based calculator to perform individualized, evidence-based risk stratification, enabling them to implement targeted prophylactic supplementation and optimized surgical planning to mitigate the risk of parathyroid injury in high-risk patients.


Introduction

Differentiated thyroid cancer (DTC), including papillary and follicular thyroid carcinomas, is the most common endocrine malignancy and accounts for the majority of thyroid cancers (1). Although DTC generally has an indolent course and a favorable prognosis, surgery remains the mainstay of treatment, particularly for patients requiring total or near-total thyroidectomy to achieve adequate oncologic control (2). With advances in operative techniques, perioperative management, and anatomical understanding, thyroid surgery has become increasingly safe (3). However, despite these improvements, postoperative complications remain clinically relevant and may adversely affect patient recovery, resource utilization, and quality of life (4).

Among these complications, postoperative hypocalcemia is the most common complication after thyroidectomy. Transient hypocalcemia has been reported in a substantial proportion of patients after total thyroidectomy, whereas permanent hypocalcemia, although less frequent, may result in persistent symptoms, long-term medication dependence, and impaired quality of life (5,6). Clinical manifestations range from mild perioral numbness and extremity tingling to tetany, laryngospasm, seizures, and cardiac arrhythmias in severe cases (7,8). Because symptoms often occur within the first 24–48 hours after surgery, postoperative hypocalcemia may lead to prolonged hospitalization, repeated biochemical monitoring, and delayed discharge, thereby increasing both patient burden and healthcare costs (9).

The pathogenesis of postoperative hypocalcemia is multifactorial. The primary mechanism is postoperative hypoparathyroidism caused by inadvertent parathyroid gland (PTG) devascularization, mechanical injury, or inadvertent excision during thyroid surgery (10). The risk may be further increased in patients undergoing more extensive procedures, such as central lymph node dissection (CLND) for malignant disease (11). Currently, perioperative parathyroid hormone (PTH) monitoring is widely used for risk assessment, but its routine use is limited by uncertainty regarding optimal measurement timing and cutoff values (12-14). More importantly, real-time PTH monitoring is costly and often unavailable in resource-limited settings (15,16). Therefore, there remains a need for practical and accessible tools to identify patients at high risk of hypocalcemia and support individualized perioperative management.

In an era characterized by rapid technological advancement and data-driven decision-making, machine learning (ML) has emerged as a transformative tool in endocrinology (17). By utilizing statistical methods to analyze complex, high-dimensional medical data, ML algorithms can uncover hidden patterns and generate highly accurate predictive models (18). In the field of thyroidectomy, ML has been used to predict postoperative outcomes, such as lymph node metastasis (19), lymph node recurrence (20), and postoperative complications (21). Nevertheless, a critical barrier to clinical adoption has been the “black-box” nature of many ML models, which limits interpretability and physician trust. An explainable artificial intelligence (XAI) framework—such as SHapley Additive exPlanations (SHAP)—addresses this limitation by quantifying individual feature contributions to model predictions, thereby enhancing transparency and supporting informed clinical decision-making (22).

In the present study, we sought to develop and validate an interpretable ML model for the early prediction of postoperative hypocalcemia in patients with DTC undergoing thyroidectomy. Using routinely available clinical variables, we aimed to construct a practical web-based tool that could facilitate individualized risk stratification and support perioperative decision-making. We present this article in accordance with the TRIPOD reporting checklist (available at https://gs.amegroups.com/article/view/10.21037/gs-2026-0211/rc).


Methods

Study design and population

This retrospective cohort study was conducted at The First Hospital of China Medical University, a tertiary referral center in Shenyang, China. A total of 1,222 patients with DTC who were treated with thyroid surgery between February 2021 and October 2025 were screened for eligibility. The inclusion criteria were as follows: (I) aged ≥18 years; (II) pathologically confirmed DTC; (III) underwent primary thyroid surgery; and (IV) no prior history of neck irradiation. The exclusion criteria were as follows: (I) abnormal baseline serum calcium levels; (II) pre-existing parathyroid dysfunction; (III) concurrent severe dysfunction of major organs (e.g., hepatic or renal insufficiency) or other malignancies; (IV) pregnancy or lactation; (V) preoperative use of medications affecting calcium metabolism (e.g., routine calcium supplements or corticosteroids); and (VI) missing data exceeding 20% for any key variables used in modeling. This study was approved by the Ethics Committee of The First Hospital of China Medical University (ethics number: EC-2025-1097-4), and was conducted in accordance with the Declaration of Helsinki and its subsequent amendments. The requirement for informed consent was waived because of the retrospective design of the study.

Definitions

Serum calcium levels were routinely measured at 8:00 AM on postoperative days (POD) 1, 2, and 3. To account for variations in serum protein concentrations, total calcium was corrected for albumin using the following formula: corrected calcium (mmol/L) = total calcium (mmol/L) + 0.02 × [40 – albumin (g/L)]. The primary endpoint, early postoperative hypocalcemia, was determined using the lowest corrected serum calcium value within the first 72 hours after surgery.

Early postoperative hypocalcemia was defined according to either of the following criteria: (I) biochemical hypocalcemia, defined as corrected serum calcium <2.1 mmol/L, regardless of the presence or absence of clinical symptoms; (II) symptomatic hypocalcemia, defined as the presence of clinical signs or symptoms suggestive of hypocalcemia despite a corrected serum calcium ≥2.1 mmol/L. Clinical manifestations included numbness or paresthesia of the fingertips, toes, or perioral region, a positive Chvostek’s sign or Trousseau’s sign, or tetany, with subsequent improvement following calcium supplementation (23).

To ensure timely detection and management of postoperative hypocalcemia, all patients underwent standardized clinical monitoring during the first 72 hours after surgery. Trained nursing staff performed physical assessments every 4–6 hours, specifically evaluating patients for symptoms and signs of hypocalcemia. Oral or intravenous calcium supplementation, with concomitant vitamin D administration when clinically indicated, was initiated immediately upon either laboratory confirmation of corrected serum calcium <2.1 mmol/L or the development of clinically significant hypocalcemic symptoms, according to institutional postoperative management protocols.

Although postoperative PTH is an established predictor of hypocalcemia, it was not routinely measured in all patients and was therefore unavailable for a substantial proportion of the cohort. Consequently, postoperative PTH was not included in the outcome definition, which was based on serial corrected serum calcium measurements and standardized clinical assessments.

Basic and clinical data collection

Baseline demographic and clinical data were extracted from the electronic medical record system. Variables collected included age, gender, body mass index (BMI), smoking history, alcohol consumption, previous non-thyroid neck surgery, comorbidities (hypertension, diabetes, and hyperlipidemia), and thyroid-related conditions (Hashimoto’s thyroiditis and Graves’ disease). General physical status was assessed using the American Society of Anesthesiologists (ASA) classification. Tumor-related variables included maximum tumor diameter, microcalcifications, multifocality, thyroid capsular invasion, lymph node metastasis, pathological T stage, and histological type (papillary or follicular thyroid carcinoma). Preoperative laboratory variables were obtained 1–3 days before surgery and included serum albumin, 25-hydroxyvitamin D [25-(OH)-VD], calcium (Ca), magnesium (Mg), PTH, thyroid-stimulating hormone (TSH), phosphorus (P), alkaline phosphatase (ALP), creatinine, blood urea nitrogen (BUN), and uric acid (UA).

Surgical variables included the extent of thyroidectomy (lobectomy or total thyroidectomy), surgical approach (conventional open, endoscopic, or robotic-assisted), CLND, lateral cervical lymph node dissection (LCLND), operation time, intraoperative bleeding loss, surgeon experience, PTG autotransplantation, and inadvertent PTG excision. Conventional open surgery was defined as transcervical thyroidectomy through a cervical incision. Endoscopic procedures primarily included transoral endoscopic thyroidectomy vestibular approach (TOETVA) and bilateral axillo-breast approach (BABA) thyroidectomy. Robotic-assisted procedures were performed using remote-access robotic approaches, including axillary and postauricular approaches.

Data preprocessing

To avoid data leakage, the dataset was randomly partitioned into a training set (70%) and a testing set (30%). Data quality was reviewed before analysis. Anomalous records—including duplicates, biologically implausible values, and data entry errors—were systematically identified and excluded. Variables were subsequently stratified into continuous and categorical variables.

To handle missing data, the extent and distribution of missingness across all features were first evaluated. Assuming that data were missing at random, missing values were imputed using multiple imputation by chained equations (MICE) (24). Details of the imputed variables are provided in Figure S1. Continuous variables were standardized using Z-score normalization, and categorical variables were transformed using one-hot encoding. All preprocessing procedures, including imputation and scaling, were performed using the training cohort only and then applied to the testing cohort. The overall study workflow is shown in Figure 1.

Figure 1 Flowchart of the study design. ROC, receiver operating characteristic; SHAP, Shapley Additive exPlanations.

Feature selection

Feature selection was performed in the training cohort only. To identify robust predictors of postoperative hypocalcemia, we used a combined feature selection strategy incorporating least absolute shrinkage and selection operator (LASSO) regression and the Boruta algorithm (25,26). For LASSO regression, the penalty parameter (λ) was selected using 10-fold cross-validation. Variables with nonzero coefficients at the one-standard-error criterion (λ1se) were retained. Boruta, a random forest-based wrapper algorithm, was employed to identify all relevant features by comparing the importance of original variables with that of randomly generated shadow features over a maximum of 500 iterations. Variables confirmed by both methods were included in the final feature set. The overlap between the two approaches was visualized using a Venn diagram.

Model development

The selected features were used to develop eight ML models in the training cohort: logistic regression (LR), random forest (RF), light gradient boosting machine (LGBM), neural network (NNet), K-nearest neighbors (KNN), multivariate adaptive regression splines (MARS), categorical boosting (CAT), and support vector machine (SVM). Model hyperparameters were tuned using grid search with 5-fold cross-validation within the training cohort. The optimal hyperparameter combination for each model was chosen according to the highest mean area under the receiver operating characteristic curve (AUC) obtained during cross-validation (details are provided in Table S1).

Model performance evaluation

Model performance was evaluated in the testing cohort. Model discrimination was assessed using the AUC with corresponding 95% confidence intervals (CIs). Receiver operating characteristic (ROC) curves were generated for each model to compare their ability to distinguish between patients with and without early postoperative hypocalcemia. To provide a more comprehensive evaluation of classification performance, confusion matrices were constructed for all models in the testing cohort. Based on the confusion matrix, additional performance metrics were calculated, including accuracy, sensitivity (recall), specificity, precision, negative predictive value (NPV), and F1-score. Calibration was evaluated using the Brier score and calibration plots, with lower Brier scores indicating better agreement between predicted probabilities and observed outcomes. In addition, Log loss was calculated to further assess the quality of probabilistic predictions, with lower values indicating better model performance.

SHAP interpretability analysis

To eliminate the “black-box” nature of the optimal ML classifier and ensure clinical transparency, SHAP was employed (27). Rooted in cooperative game theory, the SHAP framework calculates the marginal contribution of each predictive feature to the final output, providing a unified approach for both global and local model interpretability. At the macroscopic level, SHAP summary plots and feature importance plots were generated to identify the primary global drivers of early postoperative hypocalcemia across the training cohort. At the microscopic level, SHAP force plots were constructed for individual patient instances to map highly specific risk profiles, elucidating how patient-level feature interactions shifted the baseline probability of hypocalcemia toward or away from the predictive threshold.

Web-based risk calculator

To facilitate the practical application of our model in a clinical setting, the optimal ML model was deployed as an interactive web application using the Shiny framework in R software. By inputting a patient’s relevant demographic, surgical, and preoperative biochemical variables, clinicians can instantly obtain an individualized probability of early postoperative hypocalcemia for patients with DTC.

Statistical analysis

All statistical analyses were conducted using SPSS Statistics (version 27.0, IBM Corp.) and R software (version 4.4.3, R Foundation for Statistical Computing). The normality of continuous variables was assessed using the Shapiro-Wilk test. Continuous variables with a non-normal distribution are presented as the median with interquartile range (IQR) and were compared using the Mann-Whitney U test. Categorical variables are reported as frequencies and percentages, with inter-group comparisons performed using the Pearson chi-square test or Fisher’s exact test, as appropriate. All statistical tests were two-sided, and a P value <0.05 was considered statistically significant.


Results

Patient characteristics

A total of 1,222 patients with DTC who underwent thyroidectomy were included in this study. The median age was 49 years, with 747 women (61.1%) and 475 men (38.9%). Most patients were diagnosed with papillary thyroid carcinoma (93.5%). Total thyroidectomy was performed in 77.1% of patients, while conventional open surgery accounted for 58.7% of procedures. Bilateral CLND was performed in 38.9% of patients. Overall, early postoperative hypocalcemia occurred in 34.1% (417/1,222) of the cohort.

As shown in Table 1, postoperative hypocalcemia was significantly more common in female patients and those with Hashimoto’s thyroiditis or Graves’ disease. It was also more frequently observed in patients who underwent total thyroidectomy, LCLND, PTG autotransplantation, inadvertent PTG excision, or bilateral CLND, as well as in those with longer operative times. In addition, pathological T stage (T2-T3), thyroid capsular invasion, lymph node metastasis, larger maximal tumor size, lower 25-(OH)-VD, and lower Mg levels were significantly associated with postoperative hypocalcemia (all P<0.05).

Table 1

Baseline characterization between patients with or without postoperative hypocalcemia

Variables Total (N=1222) Hypocalcemia (N=417) Non-hypocalcemia (N=705) Statistics P value
Age (years) 49.00 [36.00, 62.25] 49.00 [36.00, 62.50] 49.00 [36.00, 62.50] 0.155 0.88
BMI (kg/m2) 23.50 [20.10, 27.02] 23.50 [19.95, 27.30] 23.50 [20.25, 27.00] 0.274 0.78
Gender 5.584 0.02
   Male 475 (38.9) 143 (34.3) 332 (41.2)
   Female 747 (61.1) 274 (65.7) 473 (58.8)
Smoking history 380 (31.1) 125 (30.0) 255 (31.7) 0.371 0.54
Alcohol consumption 345 (28.2) 114 (27.3) 231 (28.7) 0.250 0.62
Prior non-thyroid neck surgery 121 (9.9) 43 (10.3) 78 (9.7) 0.119 0.73
Diabetes 274 (22.4) 92 (22.1) 182 (22.6) 0.047 0.83
Hypertension 427 (34.9) 148 (35.5) 279 (34.7) 0.084 0.77
Hyperlipidemia 347 (28.4) 120 (28.8) 227 (28.2) 0.045 0.83
Hashimoto’s thyroiditis 407 (33.3) 216 (51.8) 191 (23.7) 97.453 <0.001
Graves’s disease 68 (5.6) 35 (8.4) 33 (4.1) 9.638 0.002
Surgical procedure 68.255 <0.001
   Lobectomy 280 (22.9) 38 (9.1) 242 (30.1)
   Total 942 (77.1) 379 (90.9) 563 (69.9)
Surgical approach 0.640 0.73
   Robotic 210 (17.2) 68 (16.3) 142 (17.6)
   Endoscopic 295 (24.1) 98 (23.5) 197 (24.5)
   Open 717 (58.7) 251 (60.2) 466 (57.9)
LCLND 441 (36.1) 169 (40.5) 272 (33.8) 5.408 0.02
Surgeon experience 0.070 0.79
   <100 cases/year 572 (46.8) 193 (46.3) 379 (47.1)
   ≥100 cases/year 650 (53.2) 224 (53.7) 426 (52.9)
PTG autotransplantation 626 (51.2) 231 (55.4) 395 (49.1) 4.402 0.04
Inadvertent PTG excision 430 (35.2) 231 (55.4) 199 (24.7) 113.340 <0.001
Operation time (min) 146.00 [112.75, 193.25] 152.00 [115.00, 197.50] 143.00 [111.00, 188.00] 2.045 0.04
Intraoperative bleeding (mL) 31.00 [17.00, 51.00] 32.00 [18.00, 52.00] 31.00 [17.00, 50.50] 0.251 0.80
Multifocality 378 (30.9) 132 (31.7) 246 (30.6) 0.154 0.69
CLND 149.879 <0.001
   Unilateral 747 (61.1) 156 (37.4) 591 (73.4)
   Bilateral 475 (38.9) 261 (62.6) 214 (26.6)
Pathological stage 4.606 0.03
   T1 865 (70.8) 279 (66.9) 586 (72.8)
   T2-T3 357 (29.2) 138 (33.1) 219 (27.2)
Histological type 0.029 0.86
   Papillary carcinoma 1142 (93.5) 389 (93.3) 753 (93.5)
   Follicular carcinoma 80 (6.5) 28 (6.7) 52 (6.5)
ASA score 0.066 0.97
   I 390 (31.9) 135 (32.4) 255 (31.7)
   II 512 (41.9) 174 (41.7) 338 (42.0)
   III 320 (26.2) 108 (25.9) 212 (26.3)
Thyroid capsular invasion 510 (41.7) 261 (62.6) 249 (30.9) 113.221 <0.001
Lymph node metastasis 466 (38.1) 175 (42.0) 291 (36.1) 3.940 0.047
Tumor with microcalcification 196 (16.0) 67 (16.1) 129 (16.0) 0.000 0.99
Maximal tumor diameter (cm) 1.20 [0.70, 2.30] 1.30 [0.80, 2.40] 1.20 [0.60, 2.20] 2.323 0.02
Preoperative indicators
   Albumin (g/dL) 34.00 [26.00, 42.00] 34.00 [26.00, 42.00] 34.00 [26.00, 42.00] 0.229 0.82
   25-(OH)-VD (ng/mL) 23.00 [14.00, 34.00] 16.00 [10.00, 23.00] 28.00 [18.00, 38.00] 13.639 <0.001
   Ca (mmol/L) 2.35 [2.21, 2.49] 2.36 [2.21, 2.50] 2.35 [2.22, 2.49] 0.484 0.63
   Mg (mmol/L) 23.00 [14.00, 34.00] 16.00 [10.00, 23.00] 28.00 [18.00, 38.00] 13.639 <0.001
   PTH (pmol/L) 4.50 [2.90, 6.10] 4.40 [2.80, 6.20] 4.60 [2.90, 6.10] 0.333 0.74
   TSH (μU/mL) 3.80 [2.00, 6.50] 3.80 [1.90, 7.10] 3.70 [2.00, 6.35] 0.345 0.73
   P (mmol/L) 0.92 [0.66, 1.20] 0.92 [0.67, 1.16] 0.91 [0.64, 1.21] 0.030 0.98
   ALP (U/L) 145.00 [86.00, 205.00] 150.00 [81.00, 206.00] 144.00 [87.50, 203.50] 0.378 0.71
   Creatinine (μmol/L) 52.50 [39.00, 82.00] 52.00 [39.00, 81.50] 53.00 [39.00, 82.50] 0.347 0.73
   BUN (mmol/L) 3.80 [3.00, 5.60] 3.80 [3.10, 5.90] 3.80 [3.00, 5.50] 0.718 0.47
   UA (μmol/L) 296.00 [213.00, 394.00] 299.00 [209.50, 395.00] 294.00 [215.00, 394.00] 0.148 0.88

Continuous variables are expressed as median [IQR] and compared via the Mann-Whitney U test. Categorical variables are shown as n (%) and compared using the Pearson chi-square or Fisher’s exact test. P<0.05 is statistically significant. , endoscopic procedures primarily included TOETVA and BABA thyroidectomy. 25-(OH)-VD, 25-hydroxyvitamin D; ALP, alkaline phosphatase; ASA, American Society of Anesthesiologists; BABA, bilateral axillo-breast approach; BMI, body mass index; BUN, blood urea nitrogen; Ca, calcium; CLND, central lymph node dissection; IQR, interquartile range; LCLND, lateral cervical lymph node dissection; Mg, magnesium; P, phosphorus; PTG, parathyroid gland; PTH, parathyroid hormone; TOETVA, transoral endoscopic thyroidectomy vestibular approach; TSH, thyroid-stimulating hormone; UA, uric acid.

Predictor screening and feature selection

The dataset was randomly partitioned into a training cohort (N=856) and a testing cohort (N=366) at a 7:3 ratio. Statistical analysis confirmed that there were no significant differences in baseline characteristics between the two groups (P>0.05; Table S2), indicating the comparability of the datasets for model development and validation.

To identify the most relevant predictors while accounting for feature importance, LASSO regression was utilized to perform variable selection and complexity adjustment by incorporating an L1 penalty term. Through 10-fold cross-validation, LASSO identified 11 predictive factors, including gender, Hashimoto’s thyroiditis, surgical procedure, neck dissection range, PTG autotransplantation, inadvertent PTG excision, CLND, thyroid capsular invasion, 25-(OH)-VD, Mg, and ALP (Figure 2A,2B). In parallel, the Boruta algorithm identified eight key factors demonstrating significantly higher importance than their shadow counterparts: Hashimoto’s thyroiditis, surgical procedure, inadvertent PTG excision, CLND, pathological stage, thyroid capsular invasion, preoperative 25-(OH)-VD, and serum Mg (Figure 2C). By performing a comparative analysis of the outcomes from both screening methods, a common subset of seven consensus variables was identified: Hashimoto’s thyroiditis, surgical procedure, inadvertent PTG excision, CLND, thyroid capsular invasion, preoperative 25-(OH)-VD, and serum Mg (Figure 2D). These features were ultimately utilized as the input variables for the construction of the ML predictive models.

Figure 2 Feature selection and predictor screening. (A) LASSO coefficient profiles. (B) Tuning parameter selection in the LASSO model. The left vertical dashed lines indicate the λ at the minimum cross-validation error (λmin) and the largest λ within one standard error of the minimum error (λ1se). (C) Feature importance identified by the Boruta algorithm. The ridge plot displays the importance of confirmed predictors ranked by their median Z-scores. (D) Venn diagram of consensus predictors identified by the LASSO and Boruta algorithms. 25-(OH)-VD, 25-hydroxyvitamin D; CLND, central lymph node dissection; LASSO, least absolute shrinkage and selection operator; PTG, parathyroid gland.

Notably, several baseline clinical variables, including hypertension, hyperlipidemia, diabetes mellitus, Graves’ disease, smoking history, alcohol consumption, and other demographic- and comorbidity-related factors, were initially evaluated as candidate predictors. However, these variables demonstrated limited predictive value and were not retained by the combined LASSO-Boruta feature selection process. Consequently, they were excluded from the final ML models.

Model performance and evaluation

The predictive performance of the eight optimized ML models was comprehensively evaluated in both the training and testing cohorts. During the training phase, the LGBM model demonstrated the highest discriminatory power with an AUC of 0.900 (95% CI: 0.879–0.921) (Figure 3A and Table 2). When evaluated on the unseen testing set, the LGBM model maintained superior performance with the highest AUC of 0.875 (95% CI: 0.838–0.912) (Figure 3B and Table 2). Other models also exhibited robust discrimination, with NNet and CAT achieving the second- and third-highest AUCs of 0.874 and 0.872, respectively.

Figure 3 Comparison of model performance across machine learning algorithms. (A) ROC curves in the training set. (B) ROC curves in the testing set. (C) Petal diagram of model performance in the training set. (D) Petal diagram of model performance in the testing set. AUC, area under the curve; CAT, categorical boosting; KNN, K-nearest neighbors; LGBM, light gradient boosting machine; LR, logistic regression; MARS, multivariate adaptive regression splines; NNet, neural network; NPV, negative predictive value; RF, random forest; ROC, receiver operating characteristic; SVM, support vector machine.

Table 2

Comparison of model performance matrices across the eight machine learning algorithms.

Algorithms Accuracy AUC (95% CI) Recall Precision F1-score Specificity NPV Log loss Brier score
Training cohort
   LR 0.794 0.882 (0.859–0.905) 0.548 0.784 0.645 0.922 0.798 0.435 0.138
   RF 0.791 0.868 (0.843–0.893) 0.483 0.834 0.612 0.950 0.780 0.456 0.146
   LGBM 0.828 0.900 (0.879–0.921) 0.682 0.787 0.730 0.904 0.846 0.387 0.122
   NNet 0.810 0.883 (0.860, 0.906) 0.644 0.761 0.698 0.895 0.829 0.421 0.134
   KNN 0.772 0.829 (0.801–0.857) 0.548 0.717 0.621 0.888 0.791 0.467 0.154
   MARS 0.803 0.882 (0.859–0.905) 0.620 0.757 0.682 0.897 0.820 0.410 0.131
   CAT 0.800 0.884 (0.861–0.907) 0.545 0.799 0.648 0.929 0.798 0.420 0.135
   SVM 0.808 0.883 (0.860–0.906) 0.668 0.744 0.704 0.881 0.837 0.409 0.131
Testing cohort
   LR 0.817 0.870 (0.832–0.909) 0.608 0.809 0.694 0.925 0.820 0.442 0.140
   RF 0.811 0.871 (0.832–0.910) 0.544 0.850 0.663 0.950 0.801 0.449 0.141
   LGBM 0.814 0.875 (0.838–0.912) 0.672 0.757 0.712 0.888 0.839 0.416 0.132
   NNet 0.811 0.874 (0.836–0.912) 0.672 0.750 0.709 0.884 0.839 0.429 0.136
   KNN 0.771 0.837 (0.794–0.880) 0.584 0.695 0.635 0.867 0.801 0.485 0.152
   MARS 0.828 0.863 (0.823–0.903) 0.672 0.792 0.727 0.909 0.842 0.426 0.135
   CAT 0.825 0.872 (0.835–0.910) 0.624 0.821 0.709 0.929 0.827 0.429 0.136
   SVM 0.820 0.871 (0.832–0.909) 0.728 0.740 0.734 0.867 0.860 0.424 0.134

AUC, area under the receiver operating characteristic curve; CAT, categorical boosting; CI, confidence interval; KNN, k-nearest neighbors; LGBM, light gradient boosting machine; LR, logistic regression; MARS, multivariate adaptive regression splines; NNet, neural network; NPV, negative predictive value; RF, random forest; SVM, support vector machine.

Regarding classification metrics, the LGBM model exhibited strong and balanced performance during training (Figure 3C and Table 2). In the testing set, while the MARS model achieved the highest absolute accuracy (0.828), the LGBM model maintained a highly balanced predictive profile with an accuracy of 0.814 (Figure 3D and Table 2). The RF model exhibited the highest specificity (0.950) and precision (0.850) but showed relatively low recall (0.544) and F1-score (0.663) in the testing set, indicating limited sensitivity. Notably, the SVM model demonstrated the highest sensitivity in the testing set, achieving a recall of 0.728, an NPV of 0.860, and an F1-score of 0.734 (Figure 3D and Table 2).

Calibration analysis revealed good agreement between predicted probabilities and observed outcomes across models in both the training and testing cohorts (Figure 4A,4B). In the testing set, the LGBM model showed the best calibration performance, yielding the lowest Brier score (0.132) and Log loss (0.416) (Table 2). Furthermore, as shown in Figure 5, the confusion matrices demonstrated variability in classification performance across the eight ML models in the testing cohort. The LGBM model correctly classified 214 true negatives and 84 true positives, with 27 false positives and 41 false negatives. These results indicate that the LGBM model achieved high specificity with a relatively low false-positive rate, while maintaining acceptable sensitivity in detecting patients at risk of early postoperative hypocalcemia. Therefore, based on its superior discrimination and calibration, the LGBM model was selected as the optimal model for the subsequent development of the web-based clinical prediction tool.

Figure 4 Calibration performance of the machine learning models. (A) Calibration curves in the training set. (B) Calibration curves in the testing set. CAT, categorical boosting; KNN, K-nearest neighbors; LGBM, light gradient boosting machine; LR, logistic regression; MARS, multivariate adaptive regression splines; NNet, neural network; RF, random forest; SVM, support vector machine.
Figure 5 Confusion matrices of the machine learning models in the testing set. “0” indicates the negative class, and “1” indicates the positive class. CAT, categorical boosting; KNN, K-nearest neighbors; LGBM, light gradient boosting machine; LR, logistic regression; MARS, multivariate adaptive regression splines; NNet, neural network; RF, random forest; SVM, support vector machine.

SHAP-based model interpretability

To enhance the clinical transparency of the optimal LGBM model and elucidate the primary drivers of early postoperative hypocalcemia, SHAP analysis was utilized. Figure 6A displays the features ranked in descending order of their global importance. The analysis revealed that CLND was the most critical predictor, followed by preoperative 25-(OH)-VD, thyroid capsular invasion, inadvertent PTG excision, Hashimoto’s thyroiditis, surgical procedure, and preoperative serum Mg. In the SHAP summary plot (Figure 6B), each point denotes an individual patient sample. The color gradient ranging from purple to yellow indicates the magnitude of the original feature value (from low to high). By analyzing the distribution of SHAP values, we observed that higher values of preoperative 25-(OH)-VD and serum Mg levels were associated with strong negative contributions to the risk probability. Conversely, the presence or higher values of inadvertent PTG excision, thyroid capsular invasion, Hashimoto’s thyroiditis, and more extensive surgical procedures demonstrated a strong positive influence on the model’s prediction of hypocalcemia.

Figure 6 Model interpretability based on SHAP analysis. (A) Global feature importance ranking. (B) SHAP summary plot. The color gradient represents the feature value (yellow: high; purple: low). (C) Waterfall plot for an individual high-risk patient. (D) Waterfall plot for an individual low-risk patient. Red bars indicate risk-increasing factors, while yellow bars indicate protective factors. 25-(OH)-VD, 25-hydroxyvitamin D; CLND, central lymph node dissection; LGBM, light gradient boosting machine; PTG, parathyroid gland; SHAP, Shapley Additive exPlanations.

To further enhance the understanding of the model’s decision-making process at the individual level, we conducted a detailed interpretability analysis of two representative samples using SHAP waterfall plots (Figure 6C,6D). Regarding the high-risk sample (Figure 6C), the model predicted a high hypocalcemia probability of 0.797 (baseline value: 0.342). This elevated risk was predominantly driven by the presence of bilateral CLND (+0.187), thyroid capsular invasion (+0.147), severe preoperative vitamin D deficiency (+0.116), and total thyroidectomy (+0.0676), which completely overshadowed minor protective factors such as the absence of Hashimoto’s thyroiditis (−0.0497). In contrast, in the low-risk sample (Figure 6D), the model predicted a low hypocalcemia risk probability of 0.0585. Although this patient had thyroid capsular invasion (+0.0401), this risk was successfully mitigated by strong protective factors—most notably by lobectomy (−0.131), unilateral CLND (−0.093), the absence of Hashimoto’s thyroiditis (−0.0621), and inadvertent PTG excision (−0.0381).

Development of the web-based clinical predictor

To facilitate the real-world clinical application of our findings, the optimized LGBM model was successfully deployed as an interactive, web-based risk calculator using the R Shiny framework (Figure 7). By inputting a patient’s specific perioperative parameters—including preoperative 25-(OH)-VD and Mg levels, surgical procedure details, and tumor characteristics—clinicians can use the application to instantly compute the individualized probability of early postoperative hypocalcemia. The prediction outcome is explicitly presented as a percentage risk score, allowing for rapid risk stratification. The web application is freely accessible for clinical and academic use at the following link: https://aiwebtool.shinyapps.io/earlyposthypocal/.

Figure 7 Web-based calculator based on the LGBM model for predicting early postoperative hypocalcemia in patients with DTC. 25-(OH)-VD, 25-hydroxyvitamin D; CLND, central lymph node dissection; DTC, differentiated thyroid cancer; LGBM, light gradient boosting machine; PTG, parathyroid gland.

Discussion

Early postoperative hypocalcemia is the most common complication after DTC surgery, resulting in a substantial clinical and economic burden (28). To facilitate early risk stratification and targeted prevention, we developed and validated a ML model for predicting postoperative hypocalcemia. Importantly, both total thyroidectomy and lobectomy cases were included in the cohort. Although lobectomy is associated with a lower risk of postoperative hypocalcemia due to the preservation of the contralateral PTGs, these patients represent a substantial proportion of routine clinical practice. Their inclusion enhances the model’s generalizability and real-world applicability. Moreover, surgical extent was incorporated into feature selection and model development, enabling the algorithm to account for procedure-specific differences in hypocalcemia risk. Among the eight ML algorithms evaluated, the LGBM model demonstrated superior predictive performance, achieving high discriminatory ability with an AUC of 0.875 in the testing cohort, together with excellent calibration. To enhance clinical interpretability, SHAP analysis was employed to quantify the contribution of individual predictors. This analysis identified severe preoperative 25-(OH)-VD deficiency, bilateral CLND, and thyroid capsular invasion as the most influential determinants of postoperative hypocalcemia risk. By transparently illustrating the impact of these perioperative factors, the model may facilitate individualized risk assessment and support personalized perioperative management strategies in thyroid surgery.

Several studies have previously developed predictive models for early postoperative hypocalcemia in patients with DTC. For example, Deng et al. constructed a nomogram based on a cohort of 597 post-thyroidectomy patients, demonstrating strong discriminatory performance (29). Similarly, Zhou et al. reported a nomogram for predicting postoperative hypocalcemia with good discrimination and calibration (30), while Zhu et al. developed a nomogram incorporating gender, BMI, and PTH dynamics to identify high-risk individuals (31). Despite these advances, most existing models are based on conventional logistic regression. Although such approaches offer intuitive interpretability through nomograms, their reliance on strict parametric assumptions and linear decision boundaries may restrict their ability to capture complex, non-linear relationships within high-dimensional perioperative data. To address these limitations, recent studies have explored advanced ML techniques. Muller et al. developed an ML-based model for postoperative hypocalcemia prediction with promising clinical utility (32). Likewise, Ding et al. demonstrated that both eXtreme Gradient Boosting and LR both achieved good predictive performance in patients with secondary hyperparathyroidism (33), while Rao et al. highlighted the effectiveness of artificial neural networks in this context (34). These findings collectively support the growing role of ML in improving predictive accuracy for postoperative hypocalcemia.

In the present study, we systematically evaluated eight ML algorithms in a DTC cohort, integrating discrimination and calibration to comprehensively assess model performance and clinical utility. Among the evaluated models, gradient boosting algorithms—particularly LGBM and CAT—demonstrated superior overall performance. The LGBM model achieved the highest AUC in the testing cohort and maintained a favorable balance between sensitivity and specificity, as reflected by its strong F1-score. These findings underscore the importance of gradient boosting architectures in capturing the complex, non-linear interactions inherent in postoperative hypocalcemia prediction (35). While LR achieved acceptable accuracy and provided a stable baseline, its linear structure inherently struggled to capture high-order interactions between variables—such as the compounding physiological effect of a severe preoperative 25-(OH)-VD deficiency combined with inadvertent parathyroid excision—without exhaustive manual feature engineering. Notably, the SVM model demonstrated the highest sensitivity and a strong NPV on the testing set. However, the SVM model was not selected as the final model because it is a margin-based classifier primarily designed to maximize the geometric distance between classes. This structural characteristic—optimizing hinge loss rather than a probabilistic loss function—frequently results in inferior probabilistic calibration compared with gradient-boosting frameworks, which natively optimize cross-entropy loss to yield highly calibrated risk probabilities (36-38). Ultimately, the selection of the LGBM model as the foundational algorithm for our web-based calculator was driven by its superior probabilistic calibration and balanced classification profile. By achieving the lowest Brier score and Log loss, the LGBM model ensures that the predicted risk percentages are highly reliable and closely reflect true clinical incidence, thereby assisting clinicians in optimizing personalized perioperative management in patients with DTC.

Interpreting ML models and translating their outputs into clinically meaningful insights remains a critical challenge. To enhance transparency and facilitate clinical adoption, we employed the SHAP framework to quantify feature importance and visualize the contribution of each variable to the risk of postoperative hypocalcemia. The SHAP summary analysis demonstrated that higher SHAP values correspond to a greater contribution to hypocalcemia risk, enabling clinicians to intuitively understand model predictions and move beyond “black-box” decision-making (39). This framework improves transparency and reproducibility, thereby, facilitating clinical translation. Our model identified several key predictors, including surgical extent (total thyroidectomy and bilateral CLND), tumor characteristics (thyroid capsular invasion), intraoperative factors (inadvertent PTG excision), underlying pathology (Hashimoto’s thyroiditis), and preoperative biochemical markers [25-(OH)-VD and Mg]. These findings are consistent with current knowledge regarding calcium homeostasis and surgical morbidity in DTC.

Among these factors, the extent of surgical resection emerged as a primary determinant of postoperative hypocalcemia. Total thyroidectomy and bilateral CLND, while essential for oncologic control, significantly increase the risk of parathyroid injury through devascularization, manipulation-induced stunning, or inadvertent excision (40,41). The delicate vascular supply of the PTG—particularly that of the anatomically variable inferior glands—is highly susceptible to disruption during extensive central compartment dissection. Even when the glands are preserved in situ, ischemic injury may lead to transient or permanent loss of PTH secretion (42,43). Furthermore, inadvertent PTG excision was a key mechanism underlying postoperative hypocalcemia, as the glands’ small size and resemblance to surrounding tissues make them prone to unintentional removal during thyroid surgery. The loss of PTH secretion disrupts calcium homeostasis, significantly increasing the risk of hypocalcemia (44,45). In addition, tumor capsular invasion often necessitates more aggressive surgical resection, further compounding this risk (29,30). These mechanisms collectively explain why a more extensive surgical scope is consistently associated with higher rates of postoperative hypocalcemia.

In addition to surgical factors, our model highlights the importance of underlying pathology and preoperative metabolic status. Hashimoto’s thyroiditis contributes to increased surgical difficulty due to chronic inflammation, fibrosis, and distorted anatomical planes, thereby impairing the identification and preservation of the PTGs and their blood supply. This altered microenvironment may also predispose patients to subtle preoperative disturbances in calcium-PTH dynamics, further increasing their vulnerability to postoperative hypocalcemia (31,46). Meanwhile, preoperative deficiencies in 25-(OH)-VD and Mg may critically compromise calcium homeostasis. Vitamin D deficiency reduces intestinal calcium absorption and attenuates the biological effectiveness of residual PTH activity, while hypomagnesemia impairs both PTH secretion and target-organ responsiveness, potentially leading to functional PTH resistance (47). Importantly, these metabolic abnormalities are often clinically silent yet may significantly diminish the physiological reserve required to compensate for perioperative parathyroid injury. The coexistence of these deficiencies may exert synergistic effects, further exacerbating postoperative calcium dysregulation (48,49). Collectively, these findings suggest that both surgical complexity and baseline metabolic reserve jointly determine the risk of hypocalcemia, underscoring the importance of comprehensive preoperative assessment and targeted perioperative management strategies.

Several limitations should be acknowledged. First, this was a retrospective, single-center study, which may introduce selection bias and limit the generalizability of the findings. External validation was not performed, making the model’s applicability to other geographic populations uncertain. Second, the model was constructed using clinical data collected at a single time point, which may not fully capture dynamic perioperative changes. Furthermore, the primary endpoint was limited to early postoperative hypocalcemia and did not distinguish between transient and permanent forms. Third, missing data were handled using multiple imputation, which may introduce bias despite efforts to minimize its impact. Additionally, the inclusion of intraoperative variables may reduce the model’s utility for purely preoperative risk prediction. Fourth, the study included only patients with DTC, and other thyroid malignancies were not evaluated, potentially restricting the broader applicability of the model across diverse thyroid cancer subtypes. Finally, the exclusion of patients with abnormal baseline calcium levels or pre-existing parathyroid dysfunction may limit applicability to patients with underlying calcium-parathyroid metabolic disorders. Future longitudinal studies should focus on multicenter prospective validation, the incorporation of dynamic data, and the integration of this predictive tool into routine clinical practice.


Conclusions

In summary, this study successfully developed and validated a highly accurate, interpretable ML framework to predict the risk of early postoperative hypocalcemia in patients undergoing surgery for DTC. Among the evaluated algorithms, the LGBM model demonstrated optimal predictive performance, achieving superior discrimination and excellent probabilistic calibration. By integrating SHAP interpretability analysis, we identified severe preoperative 25-(OH)-VD deficiency, bilateral CLND, and thyroid capsular invasion as the most important predictors of hypocalcemia risk. Crucially, the deployment of this optimized model as a user-friendly, freely accessible web-based calculator bridges the gap between complex artificial intelligence and practical clinical application. This transparent and highly accessible tool may assist clinicians in performing individualized, evidence-based risk stratification.


Acknowledgments

We thank the colleagues in our institution who help us collect clinical data.


Footnote

Reporting Checklist: The authors have completed the TRIPOD reporting checklist. Available at https://gs.amegroups.com/article/view/10.21037/gs-2026-0211/rc

Data Sharing Statement: Available at https://gs.amegroups.com/article/view/10.21037/gs-2026-0211/dss

Peer Review File: Available at https://gs.amegroups.com/article/view/10.21037/gs-2026-0211/prf

Funding: None.

Conflicts of Interest: Both authors have completed the ICMJE uniform disclosure form (available at https://gs.amegroups.com/article/view/10.21037/gs-2026-0211/coif). The authors have no conflicts of interest to declare.

Ethical Statement: The authors are accountable for all aspects of the work in ensuring that questions related to the accuracy or integrity of any part of the work are appropriately investigated and resolved. This study was approved by the Ethics Committee of The First Hospital of China Medical University (ethics number: EC-2025-1097-4), and was conducted in accordance with the Declaration of Helsinki and its subsequent amendments. The requirement for informed consent was waived because of the retrospective design of the study.

Open Access Statement: This is an Open Access article distributed in accordance with the Creative Commons Attribution-NonCommercial-NoDerivs 4.0 International License (CC BY-NC-ND 4.0), which permits the non-commercial replication and distribution of the article with the strict proviso that no changes or edits are made and the original work is properly cited (including links to both the formal publication through the relevant DOI and the license). See: https://creativecommons.org/licenses/by-nc-nd/4.0/.


References

  1. Ma L, Xie Y, Ke S, et al. Levothyroxine dose prediction post-thyroidectomy for differentiated thyroid carcinoma. Front Endocrinol (Lausanne) 2025;16:1727681. [Crossref] [PubMed]
  2. Brauer PR, Reddy CA, Burkey BB, et al. A National Comparison of Postoperative Outcomes in Completion Thyroidectomy and Total Thyroidectomy. Otolaryngol Head Neck Surg 2021;164:566-73. [Crossref] [PubMed]
  3. Scheller B, Culié D, Poissonnet G, et al. Recent Advances in the Surgical Management of Thyroid Cancer. Curr Oncol 2023;30:4787-804. [Crossref] [PubMed]
  4. Kim BH, Ryu SR, Lee JW, et al. Longitudinal Changes in Quality of Life Before and After Thyroidectomy in Patients With Differentiated Thyroid Cancer. J Clin Endocrinol Metab 2024;109:1505-16. [Crossref] [PubMed]
  5. Cao B, Wu G. Risk factors for hypocalcemia after total thyroidectomy: a narrative review. PeerJ 2025;13:e19808. [Crossref] [PubMed]
  6. Shahid Khan R, Iqbal H, Tahir T, et al. Postoperative Outcomes and Complication Profile Following Thyroidectomy: A Retrospective Analysis. Cureus 2025;17:e97354. [Crossref] [PubMed]
  7. Collins R, Lafford G, Ferris R, et al. Improving the Management of Post-Operative Hypocalcaemia in Thyroid Surgery. Cureus 2021;13:e15137. [Crossref] [PubMed]
  8. Kazaure HS, Zambeli-Ljepovic A, Oyekunle T, et al. Severe Hypocalcemia After Thyroidectomy: An Analysis of 7366 Patients. Ann Surg 2021;274:e1014-21. [Crossref] [PubMed]
  9. Lee YH, Liu Z, Zheng L, et al. The Rate of Postoperative Decline in Parathyroid Hormone Levels Can Predict Symptomatic Hypocalcemia Following Thyroid Cancer Surgery with Neck Lymph Node Dissection. Nutr Cancer 2025;77:1-8. [Crossref] [PubMed]
  10. Li J, Hu Z, Ma Z, et al. Reflections on prevention and treatment of post-thyroidectomy hypoparathyroidism: current management approaches and future prospects. Front Endocrinol (Lausanne) 2026;17:1768391. [Crossref] [PubMed]
  11. Chisthi MM, Nair RS, Kuttanchettiyar KG, et al. Mechanisms behind Post-Thyroidectomy Hypocalcemia: Interplay of Calcitonin, Parathormone, and Albumin-A Prospective Study. J Invest Surg 2017;30:217-25. [Crossref] [PubMed]
  12. Mathur A, Nagarajan N, Kahan S, et al. Association of Parathyroid Hormone Level With Postthyroidectomy Hypocalcemia: A Systematic Review. JAMA Surg 2018;153:69-76. [Crossref] [PubMed]
  13. Zheng LL, Chen KH, Liu ZJ, et al. The predictive value of postoperative intact parathyroid hormone for symptomatic hypocalcemia in older patients with thyroid cancer. Gland Surg 2025;14:510-9. [Crossref] [PubMed]
  14. Moreno Llorente P, García Barrasa A, Pascua Solé M, et al. Optimal cutoff values of intraoperative parathyroid hormone for predicting early and permanent hypoparathyroidism after total thyroidectomy. Langenbecks Arch Surg 2025;410:58. [Crossref] [PubMed]
  15. Patel KN, Caso R. Intraoperative Parathyroid Hormone Monitoring: Optimal Utilization. Surg Oncol Clin N Am 2016;25:91-101. [Crossref] [PubMed]
  16. Gillis A, Chen H, Wang TS, et al. Racial and Ethnic Disparities in the Diagnosis and Treatment of Thyroid Disease. J Clin Endocrinol Metab 2024;109:e1336-44. [Crossref] [PubMed]
  17. Kamińska M, Trofimiuk-Müldner M, Sokołowski G, et al. Machine learning in endocrinology: current applications and future perspectives. Endocrine 2025;90:357-66. [Crossref] [PubMed]
  18. Hong N, Park H, Rhee Y. Machine Learning Applications in Endocrinology and Metabolism Research: An Overview. Endocrinol Metab (Seoul) 2020;35:71-84. [Crossref] [PubMed]
  19. Huang J, Li Z, Zhong Q, et al. Developing and validating a multivariable machine learning model for the preoperative prediction of lateral lymph node metastasis of papillary thyroid cancer. Gland Surg 2023;12:101-9. [Crossref] [PubMed]
  20. Pang F, Wu L, Qiu J, et al. Machine learning models for diagnosing lymph node recurrence in postoperative PTC patients: a radiomic analysis. BMC Cancer 2025;25:1308. [Crossref] [PubMed]
  21. Yi Z, He E, Yang P, et al. Artificial neural network prediction of postoperative complications in papillary thyroid microcarcinoma based on preoperative ultrasonographic features. J Clin Ultrasound 2024;52:1313-20. [Crossref] [PubMed]
  22. Karim MR, Islam T, Shajalal M, et al. Explainable AI for Bioinformatics: Methods, Tools and Applications. Brief Bioinform 2023;24:bbad236. [Crossref] [PubMed]
  23. Wang X, Zhu J, Liu F, et al. Postoperative hypomagnesaemia is not associated with hypocalcemia in thyroid cancer patients undergoing total thyroidectomy plus central compartment neck dissection. Int J Surg 2017;39:192-6. [Crossref] [PubMed]
  24. White IR, Royston P, Wood AM. Multiple imputation using chained equations: Issues and guidance for practice. Stat Med 2011;30:377-99. [Crossref] [PubMed]
  25. Kursa MB, Rudnicki WR. Feature selection with the Boruta package. J Stat Softw 2010;36:1-13.
  26. Lee S, Gornitz N, Xing EP, et al. Ensembles of Lasso Screening Rules. IEEE Trans Pattern Anal Mach Intell 2018;40:2841-52. [Crossref] [PubMed]
  27. Bifarin OO. Interpretable machine learning with tree-based shapley additive explanations: Application to metabolomics datasets for binary classification. PLoS One 2023;18:e0284315. [Crossref] [PubMed]
  28. Gerardi I, Verro B, Amodei R, et al. Thyroidectomy and Its Complications: A Comprehensive Analysis. Biomedicines 2025;13:433. [Crossref] [PubMed]
  29. Deng X, Zhang N, Long K, et al. Analysis of risk factors and development of a risk prediction model for postoperative hypocalcemia in differentiated thyroid cancer. Am J Cancer Res 2025;15:3645-60. [Crossref] [PubMed]
  30. Zhou L, Tu Y, Qin S, et al. Development and validation of a nomogram for predicting postoperative hypocalcemia in patients undergoing surgery for differentiated thyroid cancer. Front Endocrinol (Lausanne) 2025;16:1628453. [Crossref] [PubMed]
  31. Zhu Q, Wang R, Xu F, et al. Prediction of early biochemical and symptomatic hypocalcemia in thyroid cancer patients after total thyroidectomy. Endocrine 2025;90:896-907. [Crossref] [PubMed]
  32. Muller O, Bauvin P, Bacoeur O, et al. Machine Learning-Based Algorithm for the Early Prediction of Postoperative Hypocalcemia Risk After Thyroidectomy. Ann Surg 2024;280:835-41. [Crossref] [PubMed]
  33. Ding C, Guo Y, Mo Q, et al. Prediction Model of Postoperative Severe Hypocalcemia in Patients with Secondary Hyperparathyroidism Based on Logistic Regression and XGBoost Algorithm. Comput Math Methods Med 2022;2022:8752826. [Crossref] [PubMed]
  34. Rao KN, Arora R, Rajguru R, et al. Artificial neural network to predict post-operative hypocalcemia following total thyroidectomy. Indian J Otolaryngol Head Neck Surg 2024;76:3094-102. [Crossref] [PubMed]
  35. Yanagawa R, Iwadoh K, Akabane M, et al. LightGBM outperforms other machine learning techniques in predicting graft failure after liver transplantation: Creation of a predictive model through large-scale analysis. Clin Transplant 2024;38:e15316. [Crossref] [PubMed]
  36. Wanyonyi M, Morris ZN, Musyoka FM, et al. Enhanced machine learning and hybrid ensemble approaches for Coronary Heart Disease prediction. PLoS One 2025;20:e0328338. [Crossref] [PubMed]
  37. Liu T, Que L, Bai W, et al. Machine learning optimization of obstructive sleep apnea screening: development and validation of a gradient boosting prediction model with a clinical implementation framework. Front Med (Lausanne) 2026;13:1775766. [Crossref] [PubMed]
  38. Ma R, Wang H, Lv C, et al. Explainable machine learning for postoperative respiratory failure prediction in open-heart surgery patients - a study based on the MIMIC-IV database. BMC Med Inform Decis Mak 2026;26:134. [Crossref] [PubMed]
  39. Mohamed YA, Khoo BE, Mohd Asaari MS, et al. Decoding the black box: Explainable AI (XAI) for cancer diagnosis, prognosis, and treatment planning-A state-of-the art systematic review. Int J Med Inform 2025;193:105689.
  40. Antakia R, Edafe O, Uttley L, et al. Effectiveness of preventative and other surgical measures on hypocalcemia following bilateral thyroid surgery: a systematic review and meta-analysis. Thyroid 2015;25:95-106. [Crossref] [PubMed]
  41. Carvalho GB, Giraldo LR, Lira RB, et al. Preoperative vitamin D deficiency is a risk factor for postoperative hypocalcemia in patients undergoing total thyroidectomy: retrospective cohort study. Sao Paulo Med J 2019;137:241-7. [Crossref] [PubMed]
  42. Do KN, Duong PT, Phung TL, et al. Predictors of Postoperative Hypocalcemia and Hypoparathyroidism Following Thyroidectomy in Hanoi, Vietnam. Int J Endocrinol Metab 2024;22:e146358. [Crossref] [PubMed]
  43. Turhan MA, Konca C, Elhan AH, et al. Incidental parathyroidectomy after thyroid surgery and relationship with postoperative hypocalcemia: a single tertiary center analysis. Updates Surg 2024;76:2573-81. [Crossref] [PubMed]
  44. Dong S, Chen Z, Shui C, et al. Risk factors for transient hypoparathyroidism and hypocalcemia following total thyroidectomy with central lymph node dissection for papillary thyroid carcinoma: a single-center retrospective study. BMC Surg 2026;26:173. [Crossref] [PubMed]
  45. Qin Y, Sun W, Wang Z, et al. A Meta-Analysis of Risk Factors for Transient and Permanent Hypocalcemia After Total Thyroidectomy. Front Oncol 2020;10:614089. [Crossref] [PubMed]
  46. Gan X, Feng J, Deng X, et al. The significance of Hashimoto’s thyroiditis for postoperative complications of thyroid surgery: a systematic review and meta-analysis. Ann R Coll Surg Engl 2021;103:223-30. [Crossref] [PubMed]
  47. Khan S, Khan AA. Hypoparathyroidism: diagnosis, management and emerging therapies. Nat Rev Endocrinol 2025;21:360-74. [Crossref] [PubMed]
  48. Tabriz N, Fried D, Uslar V, et al. Impact of Preoperative Calcium and Magnesium Supplementation on Quality of Life and Hypocalcemia Post-Thyroidectomy. Endocrinol Diabetes Metab 2026;9:e70129. [Crossref] [PubMed]
  49. Xie Y, Han F, He S, et al. Effectiveness of a multifaceted intervention on reducing non-guideline-concordant prescribing of calcium and vitamin D analogues after total thyroidectomy. Front Pharmacol 2026;17:1795773. [Crossref] [PubMed]
Cite this article as: Ye J, Wang Y. Development and validation of an explainable web-based machine learning model for predicting early postoperative hypocalcemia in patients with differentiated thyroid cancer. Gland Surg 2026;15(7):193. doi: 10.21037/gs-2026-0211

Download Citation